Frontier-hard environments for knowledge work in regulated industries
Agents are trained and evaluated primarily on software engineering and math, while the automation of knowledge work is of increasing interest to enterprises worldwide. With current models, reliable end-to-end completion of difficult tasks remains out of reach, especially in cases where correctness and rule adherence matter most.
Tasks in this category are hard to build well: realistic workspace environments demand deep domain understanding as much as knowledge of what agentic systems can do. By focusing on tasks in heavily regulated industries, we can make use of the detailed rules governing nearly all work in these companies. What restricts others becomes information-dense documentation for us.
Each task is domain-specific, grounded in the national and European law that governs the work, and deterministically verifiable. Every case has a defined correct outcome that a verifier reads back from committed system state. Building a task means getting into the domain: we pressure-test the task design with agentic QA, analyse full agent traces, probe near misses, and iterate with domain experts until each environment is fair and unambiguous.
Our first suite covers real casework under German and EU law. Domains in this sample include health insurance, housing, banking, elderly care, and trade reporting. Each task comes with detailed grounding documentation and has undergone extensive testing and multiple review stages.
We are sharing a sample of the suite. Get in touch.