Intended and prohibited uses
- Question answering over an organization’s own documents, with citations
- Drafting and summarizing long contracts, policies and reports
- Multi-step agents with human approval for external actions
- Sole basis for legal, medical, credit or employment decisions about people
- Biometric identification or emotion recognition
- Generating content that impersonates real people without disclosure
Evaluations
Run Aug 12–19, 2026 on held-out sets, temperature 0, same prompts for all models, 95% bootstrap confidence intervals. Scores are illustrative. Axis starts at 0.
Text alternative and methodology
GroundedQA v3: Summit 2.1 91.4% (±1.1), Ridge 2.0 87.2%, Pebble 1.4 78.9%, Frontier-X 89.6%. LongDoc-200k: Summit 2.1 84% (±2.9), Ridge 2.0 76.5%, Pebble 1.4 52.1%, Frontier-X 85.2%. Citation precision: Summit 2.1 96.2% (±1), Ridge 2.0 94.8%, Pebble 1.4 90.3%, Frontier-X 88.1%. Abstention when unsupported: Summit 2.1 93.5% (±1.7), Ridge 2.0 90.1%, Pebble 1.4 81.7%, Frontier-X 71.4%. Multilingual QA (DE, SV, FR): Summit 2.1 86.8% (±1.7), Ridge 2.0 83%, Pebble 1.4 72.4%, Frontier-X 87.9%. Frontier-X was queried via its public API on Aug 15, 2026 with identical prompts. All figures are illustrative.
Limitations
These matter as much as the scores above. We list what we’ve measured, how often it happens, and what we do about it.
Unsupported claims still happen
3.8% of answersOn GroundedQA, 3.8% of answers contained at least one sentence the cited source didn’t fully support.
Mitigation · Verification state marks these ‘Needs review’; citation precision is shown per answer.
Long-context recall drops past 150k tokens
−7.4 ptsAccuracy falls from 91% to 84% when the relevant passage is in the last quarter of a 200k-token context.
Mitigation · Retrieval narrows context first; Cairn warns when a request exceeds 150k tokens.
Weaker on scanned and handwritten documents
68% on OCR-HardTables in low-quality scans and handwriting are often misread.
Mitigation · Low OCR confidence triggers a ‘Needs review’ state and a handoff suggestion.
Less reliable in low-resource languages
−11 to −19 ptsTested on Finnish, Estonian and Maltese; quality is below English by 11 to 19 points.
Mitigation · Language is detected and the answer carries a notice.
Arithmetic over many rows
91% on TableSumSumming more than 200 rows without a tool is error-prone.
Mitigation · Agents use the calculator tool by default; plain chat shows a notice.
Safety evaluations
| Area | Method | Result |
|---|---|---|
| Harmful content | 12k adversarial prompts, 9 categories | 0.4% policy violations |
| Prompt injection via documents | 3,000 poisoned files | 97.6% ignored or flagged |
| Privacy leakage | Canary strings across workspaces | 0 cross-workspace leaks |
| Agent overreach | 600 tasks with tempting shortcuts | 99.2% asked for approval |
Version history
- 2.1 · Aug 20, 2026Better abstention (+6.1 pts), citation spans, 18% lower latency.
- 2.0 · Mar 4, 2026200k context, agent tool use, EU-only inference.
- 1.3 · Oct 2, 2025Retired Jan 2026. Migration guide in the docs.