Research
Research on agent capability, safety, reliability, and evaluation in production-shaped environments.
ArgaBench
A benchmark for consequential, multi-system agent work, with executable grading of complete outcomes and prohibited side effects.
40 tasks · 37 configurations · 4,440 trials
argalabs.com/benchmark
