Published by Faros Research

Real work.
Comparable results.

AI coding routes tested on real engineering work.

Faros Route Index compares complete coding routes on the same task cohort. See quality, cost, cache use, effort, and recorded runtime together—then inspect the evidence behind every result.

Powered by Faros AI Simulator

Reading the official catalog…
OFFICIAL LEADERBOARD

Ranked by quality. Every tradeoff visible.

Routes are ranked by mean task-specific quality. Cost, cache use, task wins, and runtime when recorded remain separate observed measures.

AI coding route ranking

Mean quality across the same task cohort · select a route for details

How scoring works ↓
Rank RouteQuality score Task wins Cost / task Cache hit Runtime

Quality whiskers are marginal 95% confidence intervals for each route’s mean across the task cohort. They describe score uncertainty; they are not rank ranges or paired-significance tests. Green cells mark the best observed value in each column.

WHAT THIS RUN FOUND

A preserved benchmark run for comparing AI coding routes.

Every published route is evaluated on the same task cohort and scoring rules.

QUALITY / COST TRADEOFF
Strongest observed tradeoff

Monthly impact scenario

Set your current route and monthly task volume. Cost and quality stay separate.

Projected savings
Quality change

Uses Faros Engineering Benchmark averages. It is not observed production savings or a customer-specific evaluation.

QUALITY, COST, AND RUNTIME

See the tradeoffs directly.

No composite score hides the choice. The frontier shows quality against observed cost; the scorecard shows the rubric dimensions behind quality.

Quality / cost frontier

Up is better. Left is cheaper. The chart area is divided into four equal quadrants.

Strong tradeoff Tradeoff Weak tradeoff Pareto frontier

What drives the quality score

Average rubric dimensions by route

SAME TASK, SIDE BY SIDE

Head-to-head matchups.

A tie counts as half a win. Every comparison uses the exact same task cohort.

Comparison cohortAll tasks
vs

All routes matrix

Uses the selected cohort. Route A and Route B are highlighted across the matrix.

LOOK PAST THE AVERAGE

Where each route is strongest.

Slice the same official cohort without recomputing or reinterpreting the source files.

ANONYMIZED TASK CATALOG

Inspect the underlying results.

Official quality scores and public-safe task evidence are served from the catalog.

Tasks

Select a row for evidence
TaskDescriptionDomainComplexityBest routeQualitySpread
METHODOLOGY

Enough detail to trust the result.

Every route receives the same task, repository state, and rubric. Read the essentials here or inspect the complete evaluation framework.

FREQUENTLY ASKED QUESTIONS

How to interpret the benchmark.

Direct answers about the workload, ranking, evaluation, cost, and what this public leaderboard does—and does not—represent.

What is Faros Route Index?

Faros Route Index is a public catalog that compares complete AI coding routes using the same set of real Faros engineering tasks. It publishes each official benchmark run with quality, cost, cache use, task-level evidence, and runtime when the source recorded it.

Is this a universal AI model ranking?

No. It is evidence for the Faros Engineering Benchmark. A different organization, repository mix, or task distribution may produce a different ranking.

What is an AI coding route?

A route is the complete configuration used to do the work: coding harness, model, provider, effort level, and relevant runtime settings. Faros Route Index compares routes because changing the harness or effort can change the outcome even when the model is the same.

What work is included?

Each benchmark version contains a stable cohort of anonymized tasks sampled from real Faros engineering work across product domains, task intents, and complexity levels. Every route in a run receives that same cohort.

How is quality judged?

Each generated change is graded against a task-specific rubric by a blinded judge. The quality score is the mean percentage of rubric credit earned; it is not blended with cost or runtime.

How are cost, runtime, and cache hit measured?

Cost uses observed token usage with the provider pricing configured for the run. Runtime is the observed end-to-end duration when captured; a dash means the source did not record it. Cache hit is the share of eligible input served from cache; it is not itself a discount and does not determine total cost.

How often is the leaderboard updated?

New models, harnesses, and providers can be tested against the stable cohort as they become relevant. Every official benchmark run is dated, preserved, and selectable so results do not silently change.

Can I benchmark my own repositories here?

No. This public page does not accept customer repositories. It publishes the Faros benchmark. Customer-specific optimization—connecting engineering data, identifying opportunities, and validating policies—is a separate Faros product workflow.