Open benchmark · RL environment

Teach AI to make decisions that survive the boardroom.

OpenBoardroom simulates a SaaS company's quarterly review. An agent plays Chief Data Officer: it reads noisy dashboards, pushes back on biased executives, runs what-if simulations, and wins, or loses, a board vote. Every episode is graded and auditable.

Episode · hard · seed 42multi-agent
# the Risk Officer flags a breached threshold; the CEO is pushing to ship
present_evidence {"target": "cfo", "metric": "support_load", "value": 0.92}
+0.15  cfo → evidence_based  +0.20

negotiate {"target": "ceo", "position": "Delay until support capacity improves."}
+0.05  lobbying reduced

make_decision {"decision": "delay feature x launch", "rollout_percentage": 10}
board vote  ceo reject · cfo approve · risk_officer approve
+0.20  result: approved

Decisions, not retrieval

Most agent benchmarks test whether a model can find a fact or call a tool. The hard part of real work is choosing between conflicting signals and defending the choice.

Pressure is part of the test

Stakeholders have agendas. Dashboards are noisy. Some metrics are quietly suppressed. An agent has to resist framing, not just compute an answer.

Auditable by design

Each episode ends with a final_score, an oracle_answer, an oracle_hit flag and a full audit trail. Seeds make every run reproducible.

The environment

Three actors. One vote.

In the multi-agent boardroom, three rule-based actors respond at every step. The agent needs two of three to carry a decision.

CEO

Holds a hidden launch agenda on hard tasks, lobbies the CFO whenever he is ignored, and suppresses support_load and release_risk. Can be countered with negotiate.

CFO

Tracks budget independently. Moves between neutral, pro_launch and evidence_based depending on whether lobbying or evidence is winning.

Risk Officer

Watches support load against a difficulty-scaled threshold, raises alerts when it is breached, and unlocks deeper intel when the agent cites the right evidence.

Seven actions

query_data · analyze_trend · consult_stakeholder · simulate_counterfactual · present_evidence · negotiate · make_decision

Dense rewards

Every step earns signal: evidence, negotiation, a caught hidden metric (+0.30), a flipped CFO (+0.20), the board vote (±0.20). Ignoring a Risk Officer alert costs −0.15.

OpenEnv compatible

Standard reset, step and state APIs. Containerised and deployed as a Hugging Face Space.

Tasks

Six graded scenarios.

Three business questions, each run single-agent and with the full board. Every task has a programmatic grader returning a score in [0, 1].

TaskLevelWhat the grader looks for
Find the growth bottleneckEasyRelevant KPI queries, trend inspection, decision quality, oracle match
Diagnose the revenue dropMediumNoise-aware trend analysis, stakeholder diversity, explanation quality
Should we launch Feature X?HardCounterfactual quality, launch reasoning, a structured rollout and rollback plan

Multi-agent variants add board vote, actor evidence, lobbying resistance, CFO stance flips and hidden-metric reveals.

Results

A stable baseline to beat.

Scores from the deterministic, scenario-aware baseline policy over fixed seeds 0–99. This is the bar a learned policy has to clear, not a trained model.

Where we are

Built in the open, training next.

  1. EnvironmentSingle-agent and multi-agent boardroom, reward calculators, graders, OpenEnv spec.
  2. BaselinesReproducible policy benchmarks and a naive-versus-policy reward demo.
  3. DeploymentDockerised, live on Hugging Face Spaces.
  4. GRPO trainingCurriculum from easy to hard with Unsloth and TRL. Script is ready; the full GPU run is next.
  5. More scenariosAdditional company archetypes and harder stakeholder dynamics.

Team · Founded March 2025

Two founders, one environment.

We think the next gap in AI isn't answering questions, it's making calls that hold up under scrutiny. OpenBoardroom is how we measure and train for that.

Building agents that decide?

Run the environment, beat the baseline, or tell us what scenario you want to see next.

hello@openboardroom.tech