![]()
Emergence, a frontier agentic AI lab advancing safe autonomous AI, today announced that its research arm, Emergence Research, achieved state-of-the-art results on two families of AI benchmarks: verified code generation, where an AI system must write software that is both mathematically proven correct and passes real tests, and formal reasoning, where it must construct proofs that a computer can check line by line.
This press release features multimedia. View the full release here: https://www.businesswire.com/news/home/20261006763913/en/
Emergence tops the VeriSoftBench-Aristotle benchmark with a 99% score
The results reflect the central idea behind Emergence’s approach to enterprise AI: capability alone is not enough. Before AI systems can take on consequential work, they need to understand the business context, produce a working solution, and independently confirm that the solution is correct. Emergence builds its platform, Craft, around this principle, pairing the flexibility of frontier AI models with enterprise knowledge and structured verification.
Code that is proven correct, not just plausible
VeriContest asks AI systems to solve programming problems and deliver code that passes two independent checks: a formal proof that the code does what its specification says, and a set of hidden tests that run the code and compare its output to the expected results. On the benchmark’s LeetCode problem set, Craft’s Verified Code Generation (VCG) component passed both checks on 19.16% of problems. In comparison, Claude Opus 4.7 was at 15.36%, followed by GPT-5.5 at 7.10%, Claude Sonnet 4.6 at 6.67% and Gemini 3.1 Pro at 4.78%.
The low scores across the field show how demanding the task is. Producing code that looks right is now routine for AI; producing code whose behavior can be mathematically proven and that also runs correctly is not.
Near-perfect results on formal proofs
Emergence Research also tested its VCG/Tiny Prover system on VERINA and VeriSoftBench. Both benchmarks ask AI to verify software in Lean 4, a programming language and proof assistant in which a computer checks every step of a proof. The system solved 100% of VERINA problems on the first attempt and 99% of the evaluated VeriSoftBench subset on the first attempt, reaching 100% with a second attempt.
Most AI systems today check their work by asking another AI model whether an answer looks right, which is itself a probabilistic judgment. Formal methods work differently: the AI writes a precise statement of what must be true, and a deterministic checker confirms or rejects it, with the same answer every time. This is the core of Emergence’s Neuroformal AI approach: neural models propose, reason and explore, while formal systems independently verify what must be correct.
Why this matters for enterprise AI
The biggest barrier to enterprise autonomy is less about what AI can do and more about what enterprises can safely allow it to do. An AI system that drafts an analysis can tolerate uncertainty because a person reviews the result. An AI system that writes production code, changes a manufacturing process, executes a financial transaction or modifies an enterprise system faces a different standard. The more authority an AI system receives, the more consequential an undetected error becomes.
This creates what Emergence sees as the next major challenge for enterprise AI: the verification gap. AI capabilities are advancing faster than enterprises’ ability to independently verify and govern the work those systems produce. Closing that gap requires more than better models. Consider an AI agent investigating a sudden drop in manufacturing yield. It may need to understand what the company’s data means, write a database query to retrieve it, write code to analyze it, reason about likely causes, and confirm that its conclusions and proposed actions satisfy explicit constraints. Along the way it benefits from the organization’s own history: enterprises tend to investigate similar products, processes and failures over time, and Craft’s Enterprise Knowledge capability lets agents reuse that institutional memory instead of starting from scratch.
Craft combines neural reasoning with enterprise knowledge, explicit specifications, formal verification and runtime enforcement, so that uncertainty is understood and important decisions can be independently checked before they are carried out. We call this verified autonomy: giving AI more authority to decide and act, while keeping the limits of that authority explicit, verifiable and enforceable.
These benchmark results are milestones toward that goal.
References
- N. Nguyen, A. Vempaty, A. Jagmohan, P. Dey, and R. Kokku, “SOTA on VeriSoftBench and VERINA: Interactive Lean4 Unlocks Frontier LLMs for Vericoding,” Emergence AI Blog, Sep. 2026. [Online]. Available: https://www.emergence.ai/blog/wubqc15w0squyx8mot5zn7iziw6b7x
- A. Vempaty, M. Tayal, R. Kokku, M. Allaham, and P. Dey, “VCG: A High-performance Verified Coding Agent for Python,” Emergence AI Blog, Aug. 2026. [Online]. Available: https://www.emergence.ai/blog/vcg-a-high-performance-verified-coding-agent-for-python
- Z. Ye, Z. Yan, J. He, T. Kasriel, K. Yang, and D. Song, “VERINA: Benchmarking Verifiable Code Generation,” in Proc. International Conference on Learning Representations (ICLR), 2026. arXiv:2505.23135. Available: https://arxiv.org/abs/2505.23135
- Y. Xin, Q. Chen, G. Durrett, and I. Dillig, “VeriSoftBench: Repository-Scale Formal Verification Benchmarks for Lean,” arXiv:2602.18307, 2026. Available: https://arxiv.org/abs/2602.18307
- Z. Xie, M. Pawagi, Y. Liu, A. Rai, L. Shao, J. Berberian Jr., S. Che, and W. Wang, “VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation,” arXiv:2605.08553, 2026. Available: https://arxiv.org/abs/2605.08553
Notes for Editors:
Emergence is a frontier agentic AI lab advancing safe autonomous AI through its pioneering neuroformal architecture. Its technology enables enterprises to deploy autonomous AI systems that are intelligent, verifiable and reliable enough for mission-critical environments. The Emergence team comes from the world’s leading AI labs and technology teams, including IBM Research, The Allen Institute for AI, Amazon, and Broadcom. Emergence is headquartered in New York, with offices in California, Spain, and India. Learn more at https://www.emergence.ai/
View source version on businesswire.com: https://www.businesswire.com/news/home/20261006763913/en/
Media gallery
