We recently presented a paper at ACM CAIS 2026 on safety evaluation for tool-using LLM agents. The core issue is that task completion alone can be misleading: an agent may complete a task while violating a safety or policy constraint. We separate outcomes into safe success, unsafe success, and failure, and study how verification changes this tradeoff. We evaluate this using τ-bench / Tau-bench tool-use scenarios and propose a two-tier verification architecture: deterministic policy/tool checks first, followed by an LLM-based verifier for more contextual safety cases. The main finding is that verification can reduce unsafe success, but it can also reduce task completion as the task horizon increases. This creates what we call the Verifier Tax: a horizon-dependent safety–success tradeoff in tool-using agents. Paper: https://dl.acm.org/doi/full/10.1145/3786335.3813160 Curious how others think agent evaluations should report unsafe success. Should unsafe completion be counted as success, failure, or a separate category? submitted by /u/AccomplishedLeg1508 [link] [Kommentare]
Sir Mark Rowley has asked the home secretary to introduce legislation forcing companies to publish data on stolen devices.
https://t.co/OAjdtoByRe
🖼️ ⬅️ ➡️ "Watched your Deconstruct 2019 talk and wanted to know more about the subject, so I looked up the paper. Rather ironically, something on the website that hosts it must have been tested only with English. (I have already informed them, hopefully it'll get fixed soon)" ☠️☠️☠️ via email Sep 6th 2024
Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
The kanban board where your team and AI agents work the same cards. Works with Claude, Cursor, OpenCode via MCP. Catch up any time — the board is the status update.
Understanding the team whose future lies in the hands of one seven-footer.
Hello! We launched a coin recently and my team and I really believe it is something interesting and consistent. We are not very good at marketing and there are lots of scammers out there. The owner is not used to speak in AMA’s and it can be a little hard to express his vision in VC’s. We need a spokesperson to be able to learn all that is to lear about our project and have calls with the owner to understand the vision. I am trying to help and I would really appreciate advise on where I can find someone to do this AMA’s and be the spokesperson for the owner and how much it would cost (in SOL). Any advice would be appreciated. Thank you! submitted by /u/lemon-squizzy [link] [Kommentare]
Moon, Cone, Bucket, $hroom, etc. It seems the !withdraw command does not execute. Another question: is the command !withdraw 1000 cone, or is it !withdraw 1000 cones, or do i write the token in capitals? submitted by /u/netnemirepxE [link] [Kommentare]
Everyone has been waiting for regulatory clarity before declaring crypto payments mainstream. Recent spending data shows that 53% of crypto transactions in the US are already falling into everyday categories. The argument has always been that regulation needed to come first before real adoption could follow but the interesting thing is the data suggests Americans did not wait for any of it. Will the regulatory clarity accelerate this further...I dunno you tell me source: https://www.oobit.com/news/crypto-at-the-checkout-what-americas-spending-data-reveals submitted by /u/Asleep-Equipment-593 [link] [Kommentare]