Fetching from the wire…
Research2026-09-22 · source-backed
The question safe deployment needs answered isn't whether a catastrophic trajectory can occur but how often. This method builds the importance-sampling proposal by perturbing the original model's weights, making the proposal itself a differentiably parameterized language model so the search runs by gradient descent over weight space, with an adaptive regularizer trading event amplification against estimator stability. On ~120M and ~2.6B models across three event families and 300+ rare events, it delivers over 800x compute-weighted efficiency against naive Monte Carlo below 1e-7. Implementation released. (arXiv 2609.24969)
Each link below shares sources, entities, or timing with this story.
Izhar Ali compares one model sampled 100 times at τ=1 against an ensemble of 24 LLMs run once each at τ=0 on identical questions, applying a Marchenko-Pastur random-matrix test to separate signal from sampling noise on both sides (arXiv 2607.20464). Within any single model, at...
Monte Carlo's Agent Observability runs LLM-based and rule-based evaluation monitors directly against source data in BigQuery and AWS Athena, alongside standard agent metric monitors for latency, token usage, and error rates. This fills a critical gap for teams building agents...
This comparative-statics model parameterizes the allocation between character shaping (RLHF, Constitutional AI) and rule enforcement (filters, classifiers), with closed-form expected harm plus Monte Carlo tail analysis. Optimal allocation shifts only weakly toward character sh...
arXiv 2609.18052 had Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5 solve 992 algorithmic problems as Java Spring Boot service methods against a mandated signature and DTO spec, iteration forbidden, hardcoded answers banned, producing 7,593 methods and 7,936 measured reques...
A case study tracked a repository catalog from a three-day agent-built hackathon prototype through public deployment (arXiv 2609.04711). Implementation was fast; making it trustworthy was not. The consequential problems were not crashes but plausible-but-wrong output traced to...
APort Vault replays 4,371 human-written attacks from a public CTF against a live payment agent, across 14 models from 8 labs, five policy configurations and two tracks, for 225,964 total evaluations. At Levels 2 through 4, transfers to recipients the passport didn't permit num...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.