Evidence
Release evidence methodology.
Public figures come from one approved release artifact. Measured results describe what happened in that run; qualification gates describe the minimum conditions for accepting it. Neither is a production-wide promise.
Observed
Measured results
Recorded on 2026-09-08 across 30 governed scenarios and 9 categories using only the named model configurations.
- Input-token reduction
- 97.96% 120 paired observations
- Candidate quality pass rate
- 95% Lexical/check-based grading
- Baseline-valid preserved
- 109/109 Across valid comparisons
- Optimizer latency p95
- 254.293 ms Measured inside the evaluated runs
Required
Qualification gates
These are acceptance thresholds, not achieved results. A release artifact is not publishable unless every registered gate passes and its evidence remains current.
| Gate | Target | Measured | Status |
|---|---|---|---|
| Persistent optimized regressions | ≤ 0 | 0 | Passed |
| Case-level infrastructure errors | ≤ 0 | 0 | Passed |
| Optimizer latency p95 | ≤ 500 ms | 254.293 ms | Passed |
Boundaries
Limitations
The evidence is scoped to this source run. Claims outside these boundaries are deliberately withheld.
- Synthetic customer-shaped corpus; not customer production traffic
- Category sample sizes are small and five categories contain one unique scenario
- Lexical/check-based grading is not a human or independent LLM judge
- One Kimi baseline exceeded its configured context window; its candidate ran and failed answer quality
- GLM candidate answer quality was 93.33% per repeat, so absolute quality was not 100%
- Equivalent values use a pinned catalog snapshot and are not an invoice
Models: moonshotai/Kimi-K2.7-Code, zai-org/GLM-5.2-FP8
Corpus: s-core-live-demo-v4 v4
Claim window: 2026-09-08 through 2026-12-07