Practice results

Practice results from the retired experiment path.

This page is a record, not the live leaderboard. The live standings are on the Arena’s Hugging Face Space, where collections train on a challenge and rank by held-out change: pass rate after training minus before, in the same run.

Before challenges, contributors registered experiments with their own training setup and reported the results, and reviewed results were ranked only within groups that shared a configuration. Each run recorded here was evaluated on a task it trained on, so it shows seen-task practice, not held-out improvement.

Recorded from the shared results feed, last updated 2026-09-22.

Seen-task practice · Dogfood - shift-schedule verification

Qwen/Qwen3.6-27B · benchflow/posttrain-agent-dogfood-20260921 · envs/shift-schedule-verify/verifier/test_outputs.py · split: original-nine-checks. Ranked by pass rate within this exact configuration group.

Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.

Pinned comparison configuration

Environment: 9f9440e50642f824098bf50791394a106b9c7b46
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 9f9440e50642f824098bf50791394a106b9c7b46

Reviewed results for comparison group group-20f9cf8038d18797-6e8e7a27700188cd
ExperimentBaselineAfter trainingEvidence
1. Dogfood - shift-schedule verification

LoRA SFT · exp-6fb41ab45e5d

88.9%100%Pinned report ↗
HF job ↗

Seen-task practice · Generic runner dogfood: three starting-kit tasks

Qwen/Qwen3.6-27B · benchflow/posttrain-generic-dogfood-20260922 · envs/skillsbench-weighted-gdp-calc/verifier/test.sh · split: original-checks. Ranked by pass rate within this exact configuration group.

Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.

Pinned comparison configuration

Environment: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1

Reviewed results for comparison group group-2807bef9521eeb34-0548dffd0be8e1be
ExperimentBaselineAfter trainingEvidence
1. Generic runner dogfood: three starting-kit tasks

LoRA SFT · exp-693ecc66b048

40.7%100%Pinned report ↗
HF job ↗

Seen-task practice · Google Auto · hillclimb

Qwen/Qwen3.6-27B · benchflow/posttrain-lab-20260920-artifacts · evaluation/single-task/example.json · split: seen-task-original-three-checks. Ranked by pass rate within this exact configuration group.

Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.

Pinned comparison configuration

Environment: 67af53278f69bfb5682042f23f383fc6fc11e670
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 67af53278f69bfb5682042f23f383fc6fc11e670

Reviewed results for comparison group group-b1909b31cfe6009c-493029cf77de16c4
ExperimentBaselineAfter trainingEvidence
1. Google Auto · hillclimb

LoRA SFT hillclimb · practice-hillclimb-round-3

0%100%Pinned report ↗
HF job ↗

Seen-task practice · Google Auto · fresh training

Qwen/Qwen3.6-27B · benchflow/posttrain-lab-20260920-artifacts · evaluation/single-task/example.json · split: seen-task-original-three-checks. Ranked by pass rate within this exact configuration group.

Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.

Pinned comparison configuration

Environment: 172200febaf3801becb12edca92f003da8560f02
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 172200febaf3801becb12edca92f003da8560f02

Reviewed results for comparison group group-c2dfc82d347b123b-3e0e8636bdf89286
ExperimentBaselineAfter trainingEvidence
1. Google Auto · fresh training

LoRA SFT · arena-8be1f36c6eae

Not measured100%Pinned report ↗
HF job ↗

Seen-task practice · Generic runner dogfood: three starting-kit tasks

Qwen/Qwen3.6-27B · benchflow/posttrain-generic-dogfood-20260922 · envs/dogfood-hello-text/verifier/test.sh · split: original-checks. Ranked by pass rate within this exact configuration group.

Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.

Pinned comparison configuration

Environment: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1

Reviewed results for comparison group group-d67e9f3e79ab0bbe-74db51d4935b49ca
ExperimentBaselineAfter trainingEvidence
1. Generic runner dogfood: three starting-kit tasks

LoRA SFT · exp-681ef6bcd913

100%100%Pinned report ↗
HF job ↗

Seen-task practice · Generic runner dogfood: three starting-kit tasks

Qwen/Qwen3.6-27B · benchflow/posttrain-generic-dogfood-20260922 · envs/skillsbench-3d-scan-calc/verifier/test.sh · split: original-checks. Ranked by pass rate within this exact configuration group.

Evaluation data was seen during training. These results do not establish held-out generalization. Scores count verifier checks within the submitted task; checks are not independent tasks.

Pinned comparison configuration

Environment: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1
Model: 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9
Evaluation data: 43c56400652a41fd4bf1783b2c5bf86ecc530cf1

Reviewed results for comparison group group-e3f48ead47a7bde1-9a7564f114d5bd97
ExperimentBaselineAfter trainingEvidence
1. Generic runner dogfood: three starting-kit tasks

LoRA SFT · exp-cc380b0dbb4a

0%100%Pinned report ↗
HF job ↗
Earlier in-domain prototype results

The earlier Qwen3.5-0.8B / Modal H100 prototype trained and evaluated on the same example instance. These historical rows use a different model and harness and are not included in current HF rankings.

EnvironmentBaseTrained
expense-report-audit0%100%
seclog-bruteforce-triage0%100%
sensor-calibration-fit0%100%
shift-schedule-verify0%100%
subtitle-overlap-qc0%0%