Personal Assistant Benchmark
Scores a personal assistant by what it did on the device
None defined yet.
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Scores a personal assistant by what it did on the device
A tiered, gated rubric for the Finance sector of GDPval
Document-work benchmark for healthcare persona
Evaluate AI models on journal entry audit tasks
Co-evolutionary adversarial training demo (DA vs CA)
Explore and compare RL task trajectories
RL env & benchmark for enterprise BA agents
Cached replays of 140 agent-to-agent negotiation rollouts
RL environment for sales & revenue-ops agents
RL environment & benchmark for clinical EHR agents
Run and evaluate simulated iPhone assistant tasks
Generate a personalized ad and receive a quality score
Assess AI benchmark scores with a multiβdimensional trust report
Interactive demo for the MedMosaic medical-audio benchmark