Practice results
Practice results from the retired experiment path.
This page is a record, not the live leaderboard. The live standings are on the Arena’s Hugging Face Space, where collections train on a challenge and rank by held-out change: pass rate after training minus before, in the same run.
Before challenges, contributors registered experiments with their own training setup and reported the results, and reviewed results were ranked only within groups that shared a configuration. Each run recorded here was evaluated on a task it trained on, so it shows seen-task practice, not held-out improvement.
Recorded from the shared results feed, last updated 2026-09-22.
Seen-task practice · Dogfood - shift-schedule verification
Qwen/Qwen3.6-27B · benchflow/posttrain-agent-dogfood-20260921 · envs/shift-schedule-verify/verifier/test_outputs.py · split: original-nine-checks. Ranked by pass rate within this exact configuration group.
Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.
Pinned comparison configuration
Environment: 9f9440e50642f824098bf50791394a106b9c7b46
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 9f9440e50642f824098bf50791394a106b9c7b46
| Experiment | Baseline | After training | Evidence |
|---|---|---|---|
| 1. Dogfood - shift-schedule verification LoRA SFT · exp-6fb41ab45e5d | 88.9% | 100% | Pinned report ↗ HF job ↗ |
Seen-task practice · Generic runner dogfood: three starting-kit tasks
Qwen/Qwen3.6-27B · benchflow/posttrain-generic-dogfood-20260922 · envs/skillsbench-weighted-gdp-calc/verifier/test.sh · split: original-checks. Ranked by pass rate within this exact configuration group.
Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.
Pinned comparison configuration
Environment: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
| Experiment | Baseline | After training | Evidence |
|---|---|---|---|
| 1. Generic runner dogfood: three starting-kit tasks LoRA SFT · exp-693ecc66b048 | 40.7% | 100% | Pinned report ↗ HF job ↗ |
Seen-task practice · Google Auto · hillclimb
Qwen/Qwen3.6-27B · benchflow/posttrain-lab-20260920-artifacts · evaluation/single-task/example.json · split: seen-task-original-three-checks. Ranked by pass rate within this exact configuration group.
Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.
Pinned comparison configuration
Environment: 67af53278f69bfb5682042f23f383fc6fc11e670
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 67af53278f69bfb5682042f23f383fc6fc11e670
| Experiment | Baseline | After training | Evidence |
|---|---|---|---|
| 1. Google Auto · hillclimb LoRA SFT hillclimb · practice-hillclimb-round-3 | 0% | 100% | Pinned report ↗ HF job ↗ |
Seen-task practice · Google Auto · fresh training
Qwen/Qwen3.6-27B · benchflow/posttrain-lab-20260920-artifacts · evaluation/single-task/example.json · split: seen-task-original-three-checks. Ranked by pass rate within this exact configuration group.
Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.
Pinned comparison configuration
Environment: 172200febaf3801becb12edca92f003da8560f02
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 172200febaf3801becb12edca92f003da8560f02
| Experiment | Baseline | After training | Evidence |
|---|---|---|---|
| 1. Google Auto · fresh training LoRA SFT · arena-8be1f36c6eae | Not measured | 100% | Pinned report ↗ HF job ↗ |
Seen-task practice · Generic runner dogfood: three starting-kit tasks
Qwen/Qwen3.6-27B · benchflow/posttrain-generic-dogfood-20260922 · envs/dogfood-hello-text/verifier/test.sh · split: original-checks. Ranked by pass rate within this exact configuration group.
Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.
Pinned comparison configuration
Environment: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
| Experiment | Baseline | After training | Evidence |
|---|---|---|---|
| 1. Generic runner dogfood: three starting-kit tasks LoRA SFT · exp-681ef6bcd913 | 100% | 100% | Pinned report ↗ HF job ↗ |
Seen-task practice · Generic runner dogfood: three starting-kit tasks
Qwen/Qwen3.6-27B · benchflow/posttrain-generic-dogfood-20260922 · envs/skillsbench-3d-scan-calc/verifier/test.sh · split: original-checks. Ranked by pass rate within this exact configuration group.
Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.
Pinned comparison configuration
Environment: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
| Experiment | Baseline | After training | Evidence |
|---|---|---|---|
| 1. Generic runner dogfood: three starting-kit tasks LoRA SFT · exp-cc380b0dbb4a | 0% | 100% | Pinned report ↗ HF job ↗ |
Earlier in-domain prototype results
The earlier Qwen3.5-0.8B / Modal H100 prototype trained and evaluated on the same example instance. These historical rows use a different model and harness and are not included in current HF rankings.
| Environment | Base | Trained |
|---|---|---|
| expense-report-audit | 0% | 100% |
| seclog-bruteforce-triage | 0% | 100% |
| sensor-calibration-fit | 0% | 100% |
| shift-schedule-verify | 0% | 100% |
| subtitle-overlap-qc | 0% | 0% |