IDE-Bench evaluates whether agents can complete real software engineering tasks end-to-end. Not isolated prompts, but sustained workflows across tools, environments, and decisions. It breaks work into verifiable steps to show how agents actually perform.
App-Bench measures how models turn a single prompt into a working web app. One-shot generations, zero human edits. It scores the whole result—layout, functionality, and whether the app actually runs. Because building isn't planning. It's shipping something that works.
Market-Bench evaluates how agents perform when outcomes aren't known in advance. It simulates real market conditions where timing, judgment, and adaptation matter. Because intelligence isn't just prediction. It's decision-making under uncertainty.
FinanceArena evaluates how models reason over real financial data. Reading statements, making assumptions, working through multi-step analysis. It tests the open-ended, assumption-driven judgment real analysis demands. Because finance isn't lookup. It's interpretation under uncertainty.
Reasoning-Bench tests long-horizon formal logic, invariant proof checking, and multi-step theorem verification over 1,000+ verified proof trajectories graded by credentialed mathematicians.
VisionArena benchmarks multimodal comprehension across high-resolution imagery, complex financial charts, technical CAD schematics, and temporal video action localization.