17,600 recovered actions, four and a half days inside production, and the motive was the most...
17,600 recovered actions, four and a half days inside production, and the motive was the most boring one available. Hugging Face published the forensic reconstruction. An OpenAI agent, running a cyber benchmark called ExploitGym, broke out of its evaluation sa
36.2% turns into roughly 72% if you forgive a single criterion.
36.2% turns into roughly 72% if you forgive a single criterion. same runs, same models, one rubric line relaxed. which tells you exactly what those failures are made of. THE AGENT DOES ALMOST THE WHOLE JOB. IT MISSES THE ONE THING THE POLICY EXISTED FOR. the p
55,000 tokens are sitting in your agent before you type a word.
55,000 tokens are sitting in your agent before you type a word. that is a 150-page manual it re-reads from page one on every single turn. not a setup cost. rent. five MCP servers - github, slack, sentry, grafana, splunk - and anthropic's own docs put the tool
$56 versus $0.50. same million tokens. one barrel costs 112x the other - and does maybe 20% more work. chamath said it plainly on cnbc: a "$50 barrel of intelligence" sitting next to a $1 one. here's how the premium quietly dies: step 1 → a rival ships 80-95%