U⊨22A8
new
metrics
docs
research
punchlines
U⊨22A8
⊨ Metrics
Browse metrics Public preview catalog Compare metrics Side by side
⊨ Docs
All documentation Both products, start here A8 (Touchstone) The general eval model Metrics API The catalog of named metrics
⊨ Integrate
REST API Main integration — HTTP/JSON
⊨ Research
qed-bench Benchmarks against task-appropriate baselines
punchlines
U⊨22A8 · built by @onebit0fme · Terms · Privacy

Research

Benchmarks, methodology, and the raw artifacts behind them. We publish work here when the comparisons are reproducible end-to-end and the failure modes are stateable.

  • Benchmarks May 5, 2026

    qed-bench: benchmarking metrics against task-appropriate baselines

    We trained metrics on four content-judgment tasks — holistic essay quality, SMS spam, AI-vs-human authorship, and LLM authorship attribution — and compared each one to its task-appropriate baseline: trained human raters, gold labels, or an eight-model LLM-as-judge panel. Notebooks, models, and per-judge artifacts at github.com/u22a8/qed-bench.

    Read →
← back to the landing
U⊨22A8 · built by @onebit0fme · Terms · Privacy