Blog
Latest updates from Boundless Intuition Labs.
Benchmarks, verification results, and the failures we found along the way - published as we finish a report, not on a schedule.
ResearchJul 17, 2026·12 min read
Fluency Is Not Correctness
On RuleArena's airline domain, verification lifts two frontier Claude models from 54% and 61% to 100% while cutting cost roughly fourteenfold - and a verified budget model beats both unaided frontier models.
ResearchJun 19, 2026·15 min read
A Diagnosis Should Be a Proof, Not a Probability
A frontier model gave five different diagnoses to the same patient, five times. We built a clinical classifier whose verdicts are proven in Lean 4, not sampled, and benchmarked it head-to-head.

