Semantic Gravity Fields replace the KV cache with a bounded field of meaning: a fixed footprint that keeps the whole conversation in play by merging redundancy instead of evicting it. Drop-in, no retraining, runs on a model you have not modified.
A gravity well over a semantic field: meaning accumulates and merges rather than piling up. The mark is the mechanism.
The cache scales with the length of the conversation. Capacity planning turns into guesswork the moment sessions get long.
Sliding windows and rolling summaries drop detail nobody chose to lose, usually the name, the number, the constraint.
Longer context is bought with more memory, so the answer to every roadmap question becomes another GPU.
Every tier has a stated ceiling. Nothing in the field is unbounded, so the total footprint is a number you choose before you deploy.
The gist of everything said, merged as it accumulates so the total never exceeds its budget.
Names, numbers, identifiers, and codes, held verbatim so compression never blurs them.
The latest exchange word for word, sitting over a protected instruction block that cannot be evicted.
A standard cache climbs with every turn. A field holds the ceiling you gave it and keeps answering, because redundancy is merged rather than evicted.
READ THE TECHNOLOGY →SCOPE: the figure illustrates the mechanism, not a benchmark. The O(1) line is the cache replacement, which is flat across the measured range; total deployed memory stays bounded in the low tens of MB at any depth. Measured results are reported on the evidence page with the configuration that produced them.
Relationships that survive months of conversation, not a rolling summary that forgets last week.
The memory budget is fixed before you ship. SGF is built to sit inside it.
Continuous context on milliwatts, inside a memory budget fixed before the device ships.
Accounts, cases, and prior contacts carried across threads that never reset.
Characters that remember a whole playthrough, at a cost per NPC you can budget.
Long runs that keep the plan and the constraints in view from the first commit to the last.
Bounded, inspectable memory inside the perimeter, with a footprint the auditor can be shown.
Persistent spatial and task memory on the robot, held to a footprint you set at design time.
Not on this list? If memory is your ceiling, the pilot starts with your workload. TALK TO US →
A single unbroken conversation past 1,900,000 tokens — and still running.
14.5× the model's 131,072-token trained context, with no degradation observed.
Core field footprint 19.1 MB, unchanged from 50,000 tokens through 1,900,000.
Still attends content from ~1.6 million tokens ago — the field's oldest live well, read at every generation step.
Total deployed memory ~45 MB, bounded — against ~62 GB for a standard KV cache at the same depth.
Runs on an unmodified, off-the-shelf model (Llama-3.2-1B-Instruct). No retraining, no fine-tuning.
SCOPE: the 19.1 MB field is the cache replacement, and its footprint is flat across the range above. Total deployed memory stays bounded in the low tens of MB at any depth. The compression ratio is stated at its depth and grows with conversation length — it is not a constant. This run demonstrates topical coherence and continued attention to early-conversation content at depth; it is not a claim of exact fact recall. 1.9M tokens is roughly 10% of our 20-million-token target, and no degradation was observed across it.
SCOPE: this table translates the result above into workload terms. Every published number carries its model, hardware, context length, and configuration; measurements that have not cleared that bar are not on this site.
A scoped four-week pilot: you bring one workload and a memory budget, we return a measured comparison against your current cache: footprint, recall, and the configuration that produced both.