··· ······ ··· ·······
···· ···· ····· ······
Singular is available through a controlled S-Core pilot for qualified teams. It sits in front of selected model calls, keeps the evidence your model needs, and forwards a leaner request to the provider you already use — without rewriting your application.
97.96% input-token reduction in a governed 30-scenario synthetic benchmark with live-provider calls — not a production-wide guarantee.
How it works
A gateway that optimizes context. Not a summarizer.
Singular sits between your application and your model provider. It keeps the original evidence, drops only what the model did not need, and proves the answer stayed the same — so you send fewer, leaner tokens without changing your application.
- 01
Receive
Your app sends its request to Singular exactly as it would to your provider. Nothing to rewrite — Singular speaks the OpenAI-compatible API your code already calls.
- 02
Optimize
Before the model is called, Singular selects the chunks of context that actually carry the evidence for the answer and drops the redundant noise. Repeated patterns are indexed locally and the original text is preserved — nothing is rewritten or summarized.
- 03
Forward
A smaller, evidence-dense request reaches the provider and model you choose. If anything ever looks wrong, traffic passes straight through untouched. Every decision is logged with the reason.
See the controlled loop, end to end
Follow a synthetic request through the S-Core decision and illustrative reduction, then see the provider result, Portal evidence, and confirmed rollback. Demo values are not benchmark figures or a savings guarantee.
- 01Synthetic original requestSystem policy and the active user request remain outside the marker; only the synthetic context block is eligible for reduction.
- 02S-Core selects supplied evidenceS-Core preserves policy and the request, selects the supplied accounting evidence, and drops unrelated marked context.
- 03Illustrative reductionThe demo payload changes from 1,224 to 514 tokens, an illustrative 58% reduction that is not benchmark evidence.
- 04Provider resultThe synthetic provider answer is grounded only in the retained accounting policy and supplied transaction result.
- 05Portal evidenceThe Portal shows the passed eval, approved activation gate, optimized request decision, and demo operating metrics.
- 06Workspace rollbackA confirmed workspace rollback makes the next request use the original payload before provider dispatch.
Original evidence stays intact. S-Core selects supplied context only; rollback makes the next request use the original payload.
Run it beside your traffic, measure reduction and quality on your own requests, then activate. If the numbers do not hold up, you change one URL back.
Architecture
Sits in your request path. Never in your way.
Singular is an inline, OpenAI-compatible layer between your app and your provider. It removes redundant context deterministically before the request reaches the model, keeps the evidence, and gets out of the way the moment anything looks wrong.
- Gateway
- During the controlled pilot, Singular sits in front of an agreed slice of model calls as an OpenAI-compatible API gateway. Keep the SDK, request shape, and provider you already use.
- Deterministic
- The same request always yields the same optimization. Every token we drop is logged with the reason, so results are reproducible and auditable.
- Fail-open
- If Singular ever degrades, your traffic passes straight through to the provider you selected, untouched. Your app never goes down because of us.
- Privacy-first
- We do not train on pilot prompts. Requests are encrypted in transit, and retention and access boundaries are agreed during pilot setup before traffic is enabled.
Benchmarks
Less context. Measured quality.
In 120 paired live-provider observations over a governed, synthetic 30-scenario corpus, S-Core reduced evaluated input tokens while preserving every baseline-valid comparison. The tested models were moonshotai/Kimi-K2.7-Code and zai-org/GLM-5.2-FP8. These are benchmark results, not customer production outcomes.
- 97.96%
- Input-token reductionacross governed runs
- 95%
- Candidate quality pass ratelexical/check-based grading
- 100%
- Baseline-valid preserved109 of 109 comparisons
- 254.293 ms
- Optimizer latencyp95 in measured runs
The governed result recorded 0 regressions and 0 case-level errors. Claims are limited to the named models, profile, corpus, and measurement window below.
| Model | Observations | Candidate quality | Baseline-valid preserved | Regressions | Optimizer p95 |
|---|---|---|---|---|---|
| moonshotai/Kimi-K2.7-Code | 60 | 96.67% | 100% | 0 | 254.534 ms |
| zai-org/GLM-5.2-FP8 | 60 | 93.33% | 100% | 0 | 247.131 ms |
Evidence class: live-provider · corpus class: synthetic · S-Core 0.3.0 / mvp.safe.v3 · corpus s-core-live-demo-v4 v4 · measured 2026-09-08 · claim expires 2026-12-07.
Read the evidence methodology and qualification gates →Method and limitations
- Synthetic customer-shaped corpus; not customer production traffic
- Category sample sizes are small and five categories contain one unique scenario
- Lexical/check-based grading is not a human or independent LLM judge
- One Kimi baseline exceeded its configured context window; its candidate ran and failed answer quality
- GLM candidate answer quality was 93.33% per repeat, so absolute quality was not 100%
- Equivalent values use a pinned catalog snapshot and are not an invoice
Controlled pilot
Apply for the S-Core pilot
Tell us where context cost or latency is getting in the way. We qualify one workload, establish a baseline, and only activate after the evidence and operating boundaries are agreed with your team.
Qualified applications receive a response within 48 hours.
FAQ
The questions your team will ask.
What is Singular?
Singular is a secure LLM gateway that sits in front of your model calls. It optimizes context before the request reaches your provider, so you send fewer, leaner tokens without changing your application.
How does Singular integrate with our stack?
For a qualified pilot, Singular exposes an OpenAI-compatible API gateway. We agree on one workload and a limited traffic slice; your team keeps the SDK, request shape, and provider it already calls.
Which providers and models are supported?
Pilot support is qualified against specific model configurations and workload boundaries before activation. The current public evidence names the tested Kimi and GLM configurations; broader compatibility is evaluated case by case rather than promised in advance.
Will it change the answers our model returns?
Optimization is tuned to preserve useful evidence. In the governed synthetic benchmark, candidates passed 95% of lexical/check-based quality checks and preserved all 109 baseline-valid comparisons. Those results are limited to the named models and corpus, so run Singular side by side with your own traffic before committing.
What happens if Singular is unavailable?
The gateway supports fail-open passthrough for qualified pilot routes. Timeout, fallback, and rollback behavior are verified during setup before any pilot traffic is activated.
What does Singular do with our data?
We do not train on pilot prompts. Requests are encrypted in transit, and retention, access, and support boundaries are agreed during pilot setup. Technical and security review happens before traffic is enabled.
What did the benchmark measure?
S-Core reduced evaluated input tokens by 97.96% across 120 live-provider observations on a governed 30-scenario synthetic corpus. It is not customer production traffic or a promise for every workload. Real savings depend on your traffic mix, so the honest next step is a short pilot on your own requests.
How does Singular reduce LLM costs?
Singular selects the evidence your model actually needs and drops redundant context before the request reaches your provider. Fewer input tokens can reduce long-context workload costs; the result depends on the model, pricing, and traffic mix, and must be measured on your own workload.
Can Singular optimize RAG token usage?
Yes. RAG pipelines often retrieve more context than the model needs. Singular keeps the chunks that carry the answer and drops the rest, which is where most RAG token savings come from.
Is Singular a summarizer?
No. Summarizers rewrite your context and can paraphrase or drop facts. Singular preserves the original evidence chunks and removes only what the model did not need, so the answer stays the same.
How do we get started?
Apply for the controlled pilot. We qualify one workload, agree on success and rollback criteria, run baseline and optimized traffic side by side, and activate only if the evidence holds for your requests.
Still deciding? Apply for a controlled pilot →