Open vs Proprietary IntelligenceAI Gap
AA Index v4.3 687 models Updated 10 Oct 2026
Artificial Analysis Intelligence Index

Open weights are only 4.0 months behind the proprietary frontier.

Every point is a model release. The frontier lines track the running best score in each camp, and the lag is the horizontal distance between them — how long open weights needed to reach a level proprietary models had already shipped.

Method
687 models scored
Current lag
4.0 months
Shifted curve fit, 18-month window
Best open weights
46.3
MiMo-V2.6-Pro
21 Sept 2026
Best proprietary
57.6
Claude Opus 5.5 (Max, Default Fallback)
22 Sept 2026
Latest index gap
11.3
Proprietary reached the current
open-weights level on 09 Jun 2026

The two frontiers

The stepped lines are the real frontier — each step is a model that set a new record. The soft curve behind is the smoothed trend, and the shaded band is the standing gap. Faint dots are all other releases.

Every model is scored under the current index version, v4.3, including models released years ago. When Artificial Analysis revises the index it rescores the whole catalogue, so the entire curve moves at once and older screenshots of this chart will not match. The gap between the two lines is a distance in time, so it survives a rescale far better than the scores do.

Open weights Proprietary Actual frontier Smoothed trend

Lag over time

How many months open weights trailed proprietary at each point in history, measured with the method picked above. Each point uses only what had been released by then; the dashed line is the raw gap at every open-weights record. How the methods differ.

Shifted curve fit Raw step lag

Lag by capability level

The same gap read along the capability axis instead of the calendar: to reach a given index score, how much longer did open weights need than proprietary did? Rising to the right means the gap widens as models get stronger; falling means open weights are closing in at the frontier.

Smoothed trend Actual release dates

Where the data came from

Three independent acquisition tracks are merged field by field, with the more certain source winning each field rather than one source winning outright.

TrackStatusRecords Contributes
Website crawl live 59 Open-source licence category, reasoning flag
Artificial Analysis API live 697 Intelligence Index, release dates, coverage
Leaderboard catalogue live 697 Open-weights flag for every model the API returns

Open-weights classification resolved from the website chart for 95 models, from the leaderboard catalogue for 592, and inferred for 0 where no source carried a flag.

How the lag is measured

Every method answers the same question — how many months open weights trail the proprietary frontier — but they differ in which releases they listen to. Pick one with the Method control at the top; its window sits in the menu next to it.

  • Shifted curve fit default window 12, 18 or 24 months (default 18)
    One quadratic trend is fitted to proprietary at time t and to open weights at t − lag, and the reported lag is the median over all shifts, weighted by how well each one fits. Like the parallel fit below it counts proprietary records that open weights have not matched yet, but the shared trend may bend, so it copes with the pace speeding up or slowing down. The shortest window reacts fastest and is the least settled.
  • Parallel trend fit window 12, 18 or 24 months (default 18)
    Fits one straight line through each camp’s running-best score over the window, with both lines forced to share the same slope, and reports how far apart in time they are: the vertical offset divided by the slope. Proprietary records that open weights have not matched yet lift the proprietary line, so they count the day they ship. The steadiest method when replayed month by month, but it assumes both camps improve at the same pace within the window.
  • Survival median lookback 12, 24 or 36 months (default 24)
    Treats each proprietary record in the lookback as a stopwatch that stops when an open-weights model first matches it. Records still unmatched keep running until today, and the Kaplan–Meier median of those waits is the lag. While fewer than half have been matched it can only say at least the longest wait so far. Answers “how long does catching up usually take” directly, but jumps when a single record is matched.
  • Lag as of today no parameter
    How long ago proprietary first reached the score of today’s best open-weights model. Exact and easy to check, but a lower bound for everything above that level, and it grows by a month every month until the next open-weights record — a sawtooth over time.
  • Smoothed crossing smoothing 6 to 24 months (default 12)
    The original method. Smooths both running-best curves with a local linear fit and measures the horizontal distance between them at each level open weights have reached. It cannot see past the best open-weights model, so it stays pinned to the last open-weights record and ignores newer proprietary models above it. The smoothed curves in the frontier and capability charts always use this method. The Gap column of the changelog always uses the default method, whatever is picked in the menu.