Latest data: Sep 7, 2026
Back to dashboard
Metric detail

Open vs. Closed Model Intelligence Gap

Arena AI (code leaderboard) Elo gap between the top closed and top open-weight model.

Open vs. Closed Model Intelligence Gap is at the 2th percentile since 2026, below historical norms.

Current reading

176pts

2th percentile • 0% of the way to its prior froth peak

Historical context

Red bands mark major stress windows.

What this metric is telling us now

Tracks the Elo gap, on the Arena AI code leaderboard, between the best-scoring proprietary model and the best-scoring open-weight model. A shrinking gap means open-weight releases are closing in on frontier labs' best closed models faster than those labs can extend their lead.

Why it matters

The premium valuations placed on frontier AI labs and the hyperscalers bankrolling them lean on the assumption that closed-model capability stays meaningfully ahead of free alternatives. A fast-closing gap chips away at that moat narrative even though it says nothing about near-term revenue.

Source and caveats

  • Source: Arena AI code leaderboard (via arena-ai-leaderboards mirror)
  • Update frequency: Daily (automated, keyless)
  • Last updated: Sep 7, 2026
  • Composite contribution: Context only; not included in the composite.
  • Caveats: This is a watch item, not part of the composite score. Artificial Analysis's own model-comparison API would need a $400/mo Pro plan for this data, so this instead reads github.com/oolong-tea-2026/arena-ai-leaderboards, a free, keyless, daily-scraped mirror of the Arena AI leaderboards that already tags each model's license (open vs. proprietary). That mirror is a single-maintainer side project, not Arena's own feed, and has no durability guarantee -- a prior similar mirror silently stopped updating for over a year with no visible warning. Ingestion checks the mirror's own "latest snapshot" date on every run and refuses to use data more than 4 days old; if that trips, the metric simply stops updating and a GitHub issue is filed automatically on the repo so staleness gets noticed instead of silently going unnoticed.

Methodology note

Each metric is oriented so higher means frothier, converted to a percentile against its own history, and then averaged within its category before the category scores are averaged into the composite.

This site is for educational and informational purposes only. It is not investment advice, financial advice, tax advice, or a recommendation to buy, sell, or hold any security, asset, or strategy. The metrics, the composite bubble score, and any alerts are not forecasts and are not a signal to act. Markets can stay overvalued or undervalued for long periods, and past patterns do not guarantee future results. The data is aggregated from third-party sources, is provided "as is," and may contain errors, gaps, or delays. Do your own research and consult a licensed financial professional before making any financial decision.