S1Bench live

connecting…
Read-only results · open decision models scored on the same six public suites as Jev 1.13.0, with the same prompts · refreshes every 15 s
loading…

Speed vs ability · bubble = model size

Decisions/s vs macro accuracy · each bubble is a candidate
Macro accuracy vs parameters · capability scaling

Board · sorted by macro accuracy — click a column to sort

statustargetgroup params progress macro acc Δ jev s/item dec/s ECE errs

Per-subset accuracy

Tell us what to improve · goes to a file on the bench box

Stored as plain text beside the benchmark data. No accounts, no cookies; the server keeps your message, the optional name/contact, and your IP for abuse triage only. Don’t paste anything private.

Δ compares each model with the published Jev 1.13.0 numbers on the same subsets. Bubble area ≈ log₁₀(params); a dashed ring means the parameter count is unstated. Every number is read straight from the run files, and the page is read-only.