Fetching from the wire…
Research2026-07-28 · source-backed
The first large-scale benchmark for structured ER-diagram understanding, 2,960 diagrams across curated educational sources, real-world schemas and synthetic generation, each paired with a machine-readable target schema. Common elements are fine (F1 > 0.74). Weak entities 0.28, multivalued attributes 0.14, N-ary relationships 0.07. Reasoning-augmented models gain 15-25% overall but stay sensitive to linguistic priors. If you're building schema-extraction-from-screenshot, test the hard constructs before you trust it. (arXiv 2607.24707)
Each link below shares sources, entities, or timing with this story.
Niclas Lietzow, Danielle Bitterman, and Carsten Eickhoff probe what happens when a vision-language model's eyes disagree with its memorized world knowledge, identifying a "vision-default, prior-override" causal mechanism. This is directly useful for debugging the maddening cla...
Researchers systematically evaluate whether Mamba-class state space models can replace ViT encoders in large VLMs, finding competitive performance with linear-time processing versus ViT's quadratic attention. Meaningful memory savings on high-resolution or long-context vision...
The authors formalize Dense Same-Class Attribute Misbinding and built InstaBind-Lite to measure it: 524 images, 529 groups of 3-6 same-class entities, 9,580 deterministically evaluated questions with source-instance annotations that separate copying from hallucination (arXiv 2...
TRAPSBench is a procedurally generated video benchmark of 1,404 matched physics pairs where one targeted change makes the outcome undeterminable from the visuals, scored by a Penalized Epistemic Calibration Score demanding correct answers when knowable and abstention when not....
The system improves frozen VLM agents with zero parameter updates. It collects experiences in verifiable environments, distills lessons through verifier-guided reflection, attaches a Transfer Reliability Score to each, and retrieves only relevant and reliable lessons at infere...
arXiv 2608.11816 ran 21,708 trials across nine VLMs, four elicitation paradigms, and two prompt languages. Chinese-language prompting roughly triples the odds of state-aligned framing within every model; China-origin models reframe 1.6-3.2x more than non-China models, peaking...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.