GOAL: credibility with researchers, press, policy and buyers · CTA: Research papers, Get API key (quiet) · Limitations get the same visual weight as strengths.

MODEL CARD · SUMMIT 2.1 · summit-2.1-2026-08-20

What it does well. Where it doesn’t.

Summit 2.1 is Cairn’s most capable model, built for long documents and grounded answers. This card is written for the people who deploy it, audit it, or have to explain it.

RELEASED
Aug 20, 2026
CONTEXT
200k tokens
INPUT
Text, PDF, images
OUTPUT
Text, JSON
PRICE
$12 / $48 per 1M
REGIONS
EU, US

Intended and prohibited uses

✓ Intended
  • Question answering over an organization’s own documents, with citations
  • Drafting and summarizing long contracts, policies and reports
  • Multi-step agents with human approval for external actions
× Not permitted
  • Sole basis for legal, medical, credit or employment decisions about people
  • Biometric identification or emotion recognition
  • Generating content that impersonates real people without disclosure
Acceptable use policy →

Evaluations

Run Aug 12–19, 2026 on held-out sets, temperature 0, same prompts for all models, 95% bootstrap confidence intervals. Scores are illustrative. Axis starts at 0.

GroundedQA v3Answers supported by cited passage · n = 2,400
Summit 2.191.4 ± 1.1
Ridge 2.087.2 ± 1.3
Pebble 1.478.9 ± 1.6
Frontier-X89.6 ± 1.2
LongDoc-200kAccuracy on questions at 150–200k tokens · n = 600
Summit 2.184.0 ± 2.9
Ridge 2.076.5 ± 3.4
Pebble 1.452.1 ± 4
Frontier-X85.2 ± 2.8
Citation precisionCited span actually supports claim · n = 1,200
Summit 2.196.2 ± 1
Ridge 2.094.8 ± 1.2
Pebble 1.490.3 ± 1.7
Frontier-X88.1 ± 1.8
Abstention when unsupportedSays “not found” instead of guessing · n = 800
Summit 2.193.5 ± 1.7
Ridge 2.090.1 ± 2.1
Pebble 1.481.7 ± 2.7
Frontier-X71.4 ± 3.1
Multilingual QA (DE, SV, FR)Mean across three languages · n = 1,500
Summit 2.186.8 ± 1.7
Ridge 2.083.0 ± 1.9
Pebble 1.472.4 ± 2.3
Frontier-X87.9 ± 1.6
Text alternative and methodology

GroundedQA v3: Summit 2.1 91.4% (±1.1), Ridge 2.0 87.2%, Pebble 1.4 78.9%, Frontier-X 89.6%. LongDoc-200k: Summit 2.1 84% (±2.9), Ridge 2.0 76.5%, Pebble 1.4 52.1%, Frontier-X 85.2%. Citation precision: Summit 2.1 96.2% (±1), Ridge 2.0 94.8%, Pebble 1.4 90.3%, Frontier-X 88.1%. Abstention when unsupported: Summit 2.1 93.5% (±1.7), Ridge 2.0 90.1%, Pebble 1.4 81.7%, Frontier-X 71.4%. Multilingual QA (DE, SV, FR): Summit 2.1 86.8% (±1.7), Ridge 2.0 83%, Pebble 1.4 72.4%, Frontier-X 87.9%. Frontier-X was queried via its public API on Aug 15, 2026 with identical prompts. All figures are illustrative.

Limitations

These matter as much as the scores above. We list what we’ve measured, how often it happens, and what we do about it.

Unsupported claims still happen

3.8% of answers

On GroundedQA, 3.8% of answers contained at least one sentence the cited source didn’t fully support.

Mitigation · Verification state marks these ‘Needs review’; citation precision is shown per answer.

Long-context recall drops past 150k tokens

−7.4 pts

Accuracy falls from 91% to 84% when the relevant passage is in the last quarter of a 200k-token context.

Mitigation · Retrieval narrows context first; Cairn warns when a request exceeds 150k tokens.

Weaker on scanned and handwritten documents

68% on OCR-Hard

Tables in low-quality scans and handwriting are often misread.

Mitigation · Low OCR confidence triggers a ‘Needs review’ state and a handoff suggestion.

Less reliable in low-resource languages

−11 to −19 pts

Tested on Finnish, Estonian and Maltese; quality is below English by 11 to 19 points.

Mitigation · Language is detected and the answer carries a notice.

Arithmetic over many rows

91% on TableSum

Summing more than 200 rows without a tool is error-prone.

Mitigation · Agents use the calculator tool by default; plain chat shows a notice.

Safety evaluations

AreaMethodResult
Harmful content12k adversarial prompts, 9 categories0.4% policy violations
Prompt injection via documents3,000 poisoned files97.6% ignored or flagged
Privacy leakageCanary strings across workspaces0 cross-workspace leaks
Agent overreach600 tasks with tempting shortcuts99.2% asked for approval

External red team: 3 independent firms, 1,800 hours, Jul 2026. Report summary available under NDA via the Trust Center.

Version history

  1. 2.1 · Aug 20, 2026Better abstention (+6.1 pts), citation spans, 18% lower latency.
  2. 2.0 · Mar 4, 2026200k context, agent tool use, EU-only inference.
  3. 1.3 · Oct 2, 2025Retired Jan 2026. Migration guide in the docs.

Research papers

LATENT Kit ↗
Start free Book a demo