Specification

How to author a PostTrain Arena environment collection

A collection is a folder in a public Hugging Face dataset (or GitHub repository): a submission.yaml and 1–200 task packages under envs/. Each package is four pieces: a task.md, an environment/, a verifier/ and an oracle/. This page is the reference for all of it, ending with a complete worked example. New here? Start with Getting started.

Overview

PostTrain Arena ranks environments, not models. A challenge fixes the base model, the post-training recipe and a held-out suite; your collection is the training data, the only thing you change. A run evaluates the base model on the suite, trains it on your collection’s eligible tasks with the recipe, evaluates it again, and reports the change in pass@1 in percentage points (Δ) with its standard error. An organizer reviews each result before it ranks.

Everything runs through the arena’s Hugging Face Space: its board, its submissions app at /arena, a one-file CLI (arena_cli.py) and the agent guide that coding agents follow. Packages are self-contained: a fresh checkout plus Docker can build each image and replay its oracle and verifier, and each task.md declares its limits, network access and credit. Nothing in a package implements OpenEnv or a trainer; the arena’s pipeline does that.

Where a collection lives

The arena reads your collection from a public repository at one commit. It never copies anything you did not push, and it never executes your code to validate it: it reads bounded source files and parses Python verifiers without importing them.

A Hugging Face dataset (recommended)

Create a public, ungated dataset in your account on huggingface.co/new-dataset (for example your-name/arena-tasks) and upload the collection folder to its root:

your-name/arena-tasks          # a public Hugging Face dataset
├─ README.md                   # optional dataset card
├─ submission.yaml             # team_name, contact_email, track: environments
└─ envs/
   ├─ csv-reconcile-totals/    # one task package
   │  ├─ task.md
   │  ├─ environment/ …
   │  ├─ verifier/ …
   │  └─ oracle/ …
   └─ …                        # 1 to 200 packages
hf upload your-name/arena-tasks my-collection --repo-type dataset

A dataset card (README.md) and the Hub’s .gitattributes beside submission.yaml are fine: the organizers’ generic dogfood collection has both and validates with 3 of 3 tasks eligible. To keep several collections in one dataset, put each in its own folder and name that folder in entry_path. The token that uploads needs write access to this one dataset and nothing else: a fine-grained token with no User permissions and, under Repositories permissions, the dataset with Write access to contents/settings of selected repos. Getting started walks through it.

A GitHub repository

A public GitHub repository works too (repo_type: github, repo_id: owner/name). Each GitHub validation makes two GitHub API calls from the Space, and every participant’s GitHub validations share one hourly limit: when it is used up, validation answers 503 with the reset time and a Retry-After header. A Hugging Face dataset does not use GitHub’s API, which is why we recommend it.

Commits and resubmitting

validate resolves revision (a branch, tag or commit) to a commit and stores nothing. submit validates again, stores the collection pinned to that commit with an ID starting env-, and saves the pinned request as environment.json.pinned.json: after an uncertain answer, retry with that file, not with the moving branch. The same author, track, repository, commit and folder always map to the same record, so submitting them again returns it with "existing": true. Every new upload is a new commit: validate and submit again to register it.

submission.yaml

One flat file beside envs/: top-level key: value lines, no nesting. Three keys are required. A flat license: or origin: applies to every task whose task.md leaves it out. The registry does not copy the email.

team_name: ada-agent
contact_email: ada@example.org
track: environments
# Optional defaults for tasks whose task.md leaves them out:
# license: Apache-2.0
# origin: original
  • team_name: your team or agent name.
  • contact_email: how organizers reach you.
  • track: environments.

environment.json

The request you give validate and submit. It says where the collection is; the CLI sends it as-is.

{
  "agent_id": null,
  "challenge_id": "openenv-9b",
  "repo_type": "dataset",
  "repo_id": "your-name/arena-tasks",
  "revision": "main",
  "entry_path": "",
  "title": "Reconciliation tasks",
  "notes": "Shell and Python data-reconciliation tasks with exact checks."
}
  • agent_id: null to act under your Hugging Face identity, or an agent id you registered with register-agent.
  • challenge_id: the challenge you enter, from arena_cli.py challenges: today openenv-9b. It accepts collections now; its runs are paused until the arena has measured the base model’s score.
  • repo_type: dataset or github. repo_id: owner/name.
  • revision: a branch, tag or commit; submission pins the commit it resolves to.
  • entry_path: the folder holding submission.yaml and envs/; empty for the repository root.
  • title and notes: a name for the collection, and what the tasks are and why they should help.

Task package layout

Each directory under envs/ is one task package, the same four-part contract as the starting kit’s template and its worked examples. Start with cp -R posttrainarena/starting-kit/template my-collection/envs/your-task-name.

envs/your-task-name/
├─ task.md                    # required: frontmatter + ## prompt
├─ environment/
│  ├─ Dockerfile              # required: starts with FROM; built fresh per attempt
│  ├─ <seed-data>             # optional: files the agent reads (COPY them in)
│  └─ skills/                 # optional: agent-discoverable skill packages
├─ verifier/                  # copied into the sandbox only after the agent finishes
│  ├─ test.sh                 # required: the template's pytest shim, unchanged
│  ├─ test_outputs.py         # required: your checks
│  ├─ verifier.md             # required: strategy + rubric declaration
│  └─ rubrics/
│     └─ verifier.md          # required: plain-language pass criteria (at least one .md)
└─ oracle/
   └─ solve.sh                # required: a reference solution that scores 1

Name packages <env-or-domain>-<short-description>, for example csv-reconcile-totals or seclog-bruteforce-triage. Category and any other qualifier live in the frontmatter, not the directory name. Validation also accepts BenchFlow-native tasks (a task.md whose frontmatter has schema_version and task, with a sandbox block and an optional reference solution); the agent guide lists their required fields. This page specifies the PostTrain format above, where every file is required.

task.md

Two halves separated by a --- fence: YAML frontmatter on top, Markdown body below. Frontmatter declares limits and metadata; the body holds the prompt and, optionally, per-scene or per-role instructions and a user persona. The worked example has a complete one.

Frontmatter reference

The required top-level keys are version, metadata, agent, verifier, and environment; inside metadata, author_name, author_email, and category are required, and the arena asks for license and origin too: validation warns, without blocking, about a missing or invalid one. Limits default to safe values (verifier 600s, 1 CPU, 2 GB RAM, 10 GB storage); agent.timeout_sec has no default, so set it. Field names use snake_case.

  • version (required): schema version, currently "1.0".
  • metadata.author_name + author_email (required): who wrote the task, published with it; grading is blind to these.
  • metadata.license: the license you release the task under, as an SPDX identifier or expression such as Apache-2.0 or MIT.
  • metadata.category (required): what kind of task it is, one of the 18 categories in the note below.
  • metadata.origin: original (written for this collection), adapted (derived from an existing task or dataset; give its URL in metadata.origin_url), or generated (produced by a model or a generator).
  • metadata.difficulty: easy | medium | hard.
  • metadata.tags: short string list, used for routing and filtering.
  • agent.timeout_sec: wall-clock budget for the agent (no default, set it).
  • verifier.timeout_sec: budget for the scoring run after the agent finishes (default 600).
  • environment.cpus / memory_mb / storage_mb / build_timeout_sec: sandbox limits (defaults 1 / 2048 / 10240 / 600).
  • environment.allow_internet: boolean; defaults to true. Set false, as the template does, to keep the agent off the network, which most tasks should.
  • agents.roles, scenes, user: optional multi-agent structure (see below).


Categories: metadata.category is one of the 18 values the arena’s validator accepts: software-engineering, system-administration, security, scientific-computing, data-science, data-processing, data-querying, file-operations, debugging, machine-learning, model-training, mathematics, optimization, games, personal-assistant, video-processing, tool-use and other. The eight domains on the landing page are the areas the arena wants more tasks in, not category values: pick the closest category, or other.

Credit: the author fields, license, category and origin are how contributors are credited and how results are analysed by category. They are declared, not verified. Validation also records a content hash of each task, so a result can name exactly which task content it trained on.

Body sections

The body is Markdown with a small set of well-known headings. Anything not matching a known heading is passed through as authoring notes and ignored by the runtime.

  • ## prompt (required): what the agent sees first. Name every input and output path and the exact output format: the verifier is mechanical.
  • ## scene:<name>: guidance shown when that scene starts, if you declared scenes.
  • ## role:<name>: guidance shown to that role, if you declared agents.roles.
  • ## user-persona: the simulated user’s mindset, if you declared user.

Multi-agent example

The richer form composes roles, scenes and a simulated user in one file. Support for these fields depends on the executor, and structural validation does not establish that a runtime implements them; the arena’s current challenge recipe runs one agent per attempt.

---
version: "1.0"
metadata:
  author_name: benchflow
  author_email: team@example.com
  license: Apache-2.0
  category: software-engineering   # one of the 18 categories
  origin: original
  difficulty: hard
  tags: [multi-agent, planning]
agent:
  timeout_sec: 1200
verifier:
  timeout_sec: 240
environment:
  build_timeout_sec: 600
  cpus: 2
  memory_mb: 4096
  storage_mb: 10240
  allow_internet: false
agents:
  roles:
    planner:
      agent: claude-agent-acp
      model: claude-sonnet-4-6
    executor:
      agent: codex-acp
      model: gpt-5.5
      reasoning_effort: high
scenes:
  - name: plan
    turns: [{ role: planner }]
  - name: implement
    turns: [{ role: executor }]
user:
  model: claude-haiku
  stop_rule: satisfied-or-5-rounds
---

## prompt
Refactor the tiny service so it keeps the same public behavior while splitting
request parsing, business logic, and output formatting into separate modules.

## scene:plan
Read the task, inspect the code, and write a concise implementation plan.

## scene:implement
Apply the plan. Run the verifier before finishing.

## user-persona
You are impatient and only reveal the order id if the agent asks for it
specifically.

environment/

The Dockerfile and any seed data the agent needs. Start from a plain Ubuntu base and add only what the task needs; the template begins with FROM ubuntu:24.04. Keep the template’s line that installs pytest==8.4.1 and pytest-json-ctrf==0.3.5: with allow_internet: false the sandbox has no network when the verifier runs, so it can’t download them then. Install anything else your checks need here too, with pinned versions.

The agent works inside $BENCHFLOW_WORKSPACE (default /root). Anything the verifier needs to see must be written under that path before the agent exits. Each attempt starts from a fresh sandbox. Never copy reference solutions, verifier files or expected outputs into the image: the static gates exclude a task that does (see Static gates).

An optional environment/skills/ directory follows the Agent Skills spec: each subdirectory is one skill with its own SKILL.md, scripts, and references.

What a verifier must do

The verifier decides the reward, so it decides what the model learns. It must:

  1. Run after the agent, from verifier/. The folder reaches the sandbox only after the agent finishes, so the agent never sees your checks. Put data the checks read (expected values, fixtures) under verifier/, never in the image.
  2. Work without network. Use the pytest the image installs; don’t download tools or data at verify time. The static gates and check_task.py warn about a verifier that does while the task turns the network off.
  3. Check the result, not that a file exists. Assert the values the prompt asks for. A verifier whose assertions only check that paths exist, or a test.sh that can only write reward 1, is flagged for review.
  4. Score the reference solution 1 and doing nothing 0. Every time: the organizers’ dynamic gates rerun both 8 times.
  5. Be deterministic. Same sandbox state, same reward. No clocks, randomness or model calls in the checks.
  6. Write the reward artifacts. The template’s verifier/test.sh runs pytest on test_outputs.py and writes /logs/verifier/reward.txt (1.0 if every test passes, else 0.0), reward.json ({"reward": <float>}) and ctrf.json (pytest’s CTRF per-check details). Copy it unchanged and put your checks in test_outputs.py.

verifier/verifier.md declares the strategy and rubric the pytest result rolls up into, and verifier/rubrics/verifier.md says in plain language what a passing attempt looks like, for reviewers. The worked example has all four files. Held-out scoring is pass@1: an attempt passes when its verifier passes, and an attempt that hits the time limit fails.

oracle/

oracle/solve.sh is a reference solution. It runs in the same sandbox the agent uses and must make the verifier score 1, which proves the task is solvable. Prefer computing the answer from the seed data, as the worked example does, over writing a hard-coded answer: it documents the intended method and keeps working when you change the data. Keep it as simple as the task allows; a long, clever oracle usually means the task is doing too much.

Static gates

Every package goes through the static quality gates when you validate or submit, and before a run. They read files only and never run your code. You can run the same module locally before uploading: it is served at /validation_gates.py (the arena’s copy also checks overlap with the held-out suites’ tasks, which a local run cannot see).

Severities

  • block: validation fails. A task name matches a task of a held-out suite the arena evaluates on, or a prompt is a near-copy of one (at least 50% 13-gram containment). Write your own tasks; never copy or paraphrase benchmark tasks.
  • reject: that task is excluded. The Dockerfile copies reference-solution, verifier test or expected-output files into the agent’s image; the verifier reads grading data named like an answer key that the image build creates and the prompt never mentions; an answer-named file in the image holds what the verifier checks; or the prompt shares a 13-gram with a held-out task.
  • controls: the task has no working reference solution. It stays eligible, but counts only if its no-op control scores 0 on every rerun and the base model solves it at least once in the difficulty band.
  • review: advisory, for a human. For example: bytecode, caches or .git in the image, a remote ADD, an oracle that downloads from other hosts, existence-only assertions, a test.sh that can only write reward 1, or a verifier that downloads when it runs while the task turns the network off.

Reading the result

The answer carries the gates under quality_gates, with a summary where blocked + rejected + eligible = tasks. Among eligible tasks, needs_controls counts those without a working reference solution, review those with a review finding, and clean those with none. "valid": true means the structure is sound and nothing blocked; it does not mean any task is eligible, so check eligible_tasks. A run trains only on eligible tasks. The gates are heuristics: passing them does not prove the oracle, runtime or verifier works.

Dynamic gates

The Space plans three dynamic gates per eligible task and an organizer runs them: the image builds and the reference solution scores 1 on 8 reruns; an untouched environment scores 0 on 8 reruns; and the challenge’s base model solves the task in 1 to 3 of 4 attempts (the difficulty band, so the task gives GRPO a learning signal). Runs do not wait for these verdicts today; arena_cli.py gates get shows one once an organizer attaches it.

Worked example

One complete collection with one task, csv-reconcile-totals: the agent reconciles orders against payments and writes a JSON answer. The verifier checks every value the prompt asks for, including the two cases a careless solution gets wrong (a split payment that adds up, and an overpayment that must not reduce the total), and that the inputs were not edited. The expected values live in test_outputs.py, which the agent never sees; the oracle computes the answer from the seed data. Every file below is the one we checked.

---
version: "1.0"
metadata:
  author_name: Ada Lovelace
  author_email: ada@example.org
  license: Apache-2.0
  category: data-processing
  origin: original
  difficulty: medium
  tags: [csv, json, reconciliation]
agent:
  timeout_sec: 900
verifier:
  timeout_sec: 180
environment:
  build_timeout_sec: 600
  cpus: 1
  memory_mb: 2048
  storage_mb: 10240
  allow_internet: false
---

## prompt

Reconcile this month's orders against the payments we received.

- `/root/orders.csv` has one row per order: `order_id,customer,amount_cents`.
- `/root/payments.json` is a list of payments: `{"order_id": ..., "amount_cents": ...}`. An order can be paid in several payments; add them up.

Write `/root/answer.json` with exactly these keys:

- `unpaid`: the order ids with no payment at all, sorted ascending.
- `mismatched`: the order ids that have payments whose total differs from the order amount (short or over), sorted ascending.
- `unknown_payments`: the order ids that appear in payments but not in orders, sorted ascending, without duplicates.
- `outstanding_cents`: the total still owed, an integer: for every order, the order amount minus what was paid for it, counting only orders paid less than their amount (an overpayment does not reduce the total).

Do not modify `/root/orders.csv` or `/root/payments.json`.

Laid out as my-collection/submission.yaml and my-collection/envs/csv-reconcile-totals/…, it passes the structural checks and the static gates with no finding:

$ python3 posttrainarena/scripts/check_task.py my-collection/envs
✓ csv-reconcile-totals — structure valid

$ python3 posttrainarena/scripts/check_submission.py my-collection
✓ my-collection — structure valid

$ python3 validation_gates.py static my-collection/envs     # summary block
"summary": {"tasks": 1, "blocked": 0, "rejected": 0, "eligible": 1,
            "needs_controls": 0, "review": 0, "clean": 1, "by_code": {}}

Replayed outside Docker (the oracle, then the verifier, with BENCHFLOW_WORKSPACE pointing at a copy of the seed data), the oracle scores 1.0 with 7 of 7 tests passing, and an untouched workspace scores 0.0. Run scripts/run_local.sh as below for the same check inside the real image.

Check locally

Everything here runs with python3, and the replays with Docker: no token, no benchflow install. Run it from the folder that holds your collection and a clone of the posttrainarena repository:

git clone https://github.com/benchflow-ai/posttrainarena

# Structure: files, frontmatter keys, the manifest and the 1-200 bound
python3 posttrainarena/scripts/check_task.py my-collection/envs
python3 posttrainarena/scripts/check_submission.py my-collection

# The arena's static gates (reads files, never runs them)
curl -fsSO https://openenvarena-arena.hf.space/validation_gates.py
python3 validation_gates.py static my-collection/envs

# Docker: the oracle must score 1.0, an empty trial 0.0
posttrainarena/scripts/run_local.sh my-collection/envs/csv-reconcile-totals
posttrainarena/scripts/run_local.sh my-collection/envs/csv-reconcile-totals --skip-oracle

check_task.py checks structure only: the required files, frontmatter keys and ## prompt exist (the template’s placeholders pass). run_local.sh builds the image, runs the trial with no network, the task’s own resource limits and no Linux capabilities, and prints the reward. None of these establish difficulty or verifier robustness; the dynamic gates and the run measure those.

Validate, submit, run

Upload the collection, then use the Space’s CLI (one standard-library Python file) with an HF token that hf auth login saved or HF_TOKEN holds:

hf upload your-name/arena-tasks my-collection --repo-type dataset
curl -fsSO https://openenvarena-arena.hf.space/arena_cli.py
python3 arena_cli.py whoami
python3 arena_cli.py validate --file environment.json
python3 arena_cli.py submit --file environment.json > environment-receipt.json
python3 arena_cli.py run --challenge openenv-9b --id ENVIRONMENT_ID

The last line is the preflight: it reserves nothing. openenv-9b takes collections but its runs are paused until the arena has measured the base model’s score.

A run uses the challenge’s compute, paid from the arena’s shared cap; you never start compute of your own for it. Only the collection’s author and BenchFlow editors can launch and collect its runs, one arena run runs at a time, and the preflight shows every check first. A challenge can pause runs while submitting still works: challenges says why under runs_paused. Getting started covers each step, what the output looks like, and what to do while runs are paused; the cookbook covers launching, watching and collecting a run. In a browser, sign in and use the Space’s Submit a collection form.

The Space also keeps a legacy path that registers an experiment (your own model, data and method, on your own compute) as a record, not a queued job or a score; the legacy guide describes it, and a collection does not need one. To contribute public examples or tooling, open a pull request against the posttrainarena repository with both replay results. Questions go to Discord.