feat(routing): automatic prompt caching + a per-session routing budget (#982)

* feat(routing): cache-aware routing and zero-config prompt caching

Reconcile prompt caching with intelligent routing: zero-config implicit
session affinity, cache_control preservation/injection across transforms,
cache-adjusted routing economics, and a response-driven cache-hit feedback loop.

* refactor(caching): model-aware cache markers, drop cache-aware routing

* feat(routing): session stickiness cache-regret cost gate

* refactor(routing): expose per-model input and cached rates

* fix(docker): patch libssh2 CVEs by upgrading after package install

* feat: record counterfactual route on vetoed session switches

Add opt-in session_stickiness.record_counterfactual that emits the
plano.switch.counterfactual_route OTEL attribute when the cache-regret
gate retains the previous model, capturing the route it would have taken
had the switch been allowed. Telemetry only; the candidate is never
dispatched. Includes schema, demo config, and a getting-started guide.

* refactor: group prompt-caching and session-stickiness logic into cohesive modules

Extract the two concerns interleaved in the LLM handler into dedicated
sibling modules for readability: prompt_caching.rs (cache-marker injection)
and session_stickiness.rs (session key/prefix-hash resolution, pin lookup,
cache-regret gate, and pin planning). handlers/llm/mod.rs is now a thin
orchestrator that calls them at clear phase boundaries. Behavior-preserving.

* refactor(routing): unify session routing into a cumulative switch budget

Replace observed_cache_hit + per-request threshold with time/provider-TTL
warmth and a per-session switch budget in a single session_router::route()
used by both the full-proxy and /routing decision paths.

* refactor(routing): move switch budget to routing.routing_budget and rename telemetry

* refactor(routing): drop negative-cost budget refund from routing_budget

* refactor(routing): express routing_budget as a percentage overhead cap

Replace the absolute seed_usd pool with max_overhead_pct (whole-number
percent) measured against the session's running never-switch baseline:
allow a paid switch only while cumulative switch spend stays within
max_overhead_pct% of what staying on the anchor would have cost. Track
baseline_usd and switch_spend_usd on the binding (monotonic, no refunds).

* refactor(observability): align routing-budget telemetry with the overhead cap

Rename the switch-decision reasons (within_budget/over_budget -> within_cap/
over_cap) and the session span attributes to match the percentage overhead-cap
model: budget_remaining_in_usd -> overhead_pct + switch_spend_in_usd +
baseline_in_usd, and switch threshold_in_usd -> overhead_ceiling_in_usd.

* feat(observability): emit per-request and cumulative session cost

Price each turn from the catalog rates and surface it as llm.usage.{input,
output,total}_cost_usd on the (llm) span, then accumulate into a conversation-
level session_cost_usd on the binding and emit plano.session.total_cost_in_usd
on the routing span. Cache-creation tokens are priced at the plain input rate,
and the OpenAI vs Anthropic prompt-token convention (cached folded in vs
reported separately) is captured at parse time so the cost math is correct for
both. Rates are resolved request-side since the response path is synchronous.

* refactor(routing): price the never-switch baseline against the session default model

Distinguish the session's default_model (the model it started on, i.e. what it
would have cost by never switching) from anchor_model (the model that handled
the latest request, which the session is warm on). The never-switch baseline now
grows at the default model's cached rate rather than the drifting anchor's, so
the overhead-cap denominator stays true to "% above never-switching"; switch cost
is still measured against the current anchor. Also rename estimate_context_tokens
-> actual_context_tokens (and est_context_tokens -> context_tokens) since the
router prefers the provider's real prompt-token count when the session is warm.

* feat(routing): track bounded route history and price returns to still-warm models

Record a bounded (LRU, capped) per-model visit history on the session binding
and use it to sharpen the switch-cost estimate: a return to a model still within
its cache window re-reads only the tokens appended since its last visit (at its
cached rate) instead of the whole context at the uncached rate, so an A->B->A
switch is no longer over-charged as a full re-ingest. History is carried through
the response-side refresh (refining the anchor's entry from real usage) and
cleared on prefix drift. Adds the plano.switch.candidate_warm_tokens span
attribute.

* refactor(routing): tidy test names and rename pin events to binding events

Shorten run-on test names and replace stale "pin" vocabulary in the session
binding metric (brightstaff_session_pin_events_total -> _binding_events_total),
dropping lifecycle event labels that are no longer emitted.

* feat(routing): price switch cost against uncached rates when caching is off

The overhead-cap gate previously always priced the anchor at its cached rate,
assuming a warm provider cache. When prompt caching is disabled there is no
warm cache to lose, so both the switch cost and the never-switch baseline are
now priced at the plain uncached input rate:
switch_cost = context_tokens x (candidate_uncached - anchor_uncached) / 1M.

* fix(brightstaff): bound pricing-feed fetch with timeouts and warn when session state runs without a tenant header

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(routing): validate the session overhead cap end-to-end across a multi-turn session

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(demos): fix stale pinning semantics and add routing_budget demo script

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Adil Hafeez <adil.hafeez@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Musa 2026-07-20 17:53:34 -07:00 committed by GitHub
parent 80bb044857
commit 844f08bda7
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
38 changed files with 4639 additions and 316 deletions

View file

@ -273,6 +273,13 @@ static_resources:
auto_host_rewrite: true
cluster: {{ cluster_name }}
timeout: 300s
{% if cluster.prefix_affinity %}
# Feeds the cluster's RING_HASH lb_policy: requests with the
# same prompt-prefix hash stick to the same replica.
hash_policy:
- header:
header_name: x-plano-prefix-hash
{% endif %}
{% endfor %}
http_filters:
- name: envoy.filters.http.router
@ -995,9 +1002,18 @@ static_resources:
{% else -%}
connect_timeout: {{ upstream_connect_timeout | default('5s') }}
{% endif -%}
{% if cluster.prefix_affinity -%}
# KV-aware replica stickiness: resolve every replica address and consistent-hash
# requests (by x-plano-prefix-hash, see the route-level hash_policy) so the same
# prompt prefix lands on the replica holding its warm KV cache.
type: STRICT_DNS
dns_lookup_family: V4_ONLY
lb_policy: RING_HASH
{% else -%}
type: LOGICAL_DNS
dns_lookup_family: V4_ONLY
lb_policy: ROUND_ROBIN
{% endif -%}
load_assignment:
cluster_name: {{ cluster_name }}
endpoints:

View file

@ -155,6 +155,13 @@ properties:
- https
http_host:
type: string
prefix_affinity:
type: boolean
description: >
For self-hosted multi-replica backends (e.g. vLLM): consistent-hash requests
to replicas by the x-plano-prefix-hash header so the same prompt prefix lands
on the same replica and reuses its KV cache. Uses ring-hash load balancing
across all resolved endpoint addresses. Default false.
additionalProperties: false
required:
- endpoint
@ -316,6 +323,33 @@ properties:
orchestrator_model_context_length:
type: integer
description: "Maximum token length for the orchestrator/routing model context window. Default is 8192."
prompt_caching:
type: object
description: >
Automatic provider prompt caching, configured once for the whole Plano instance.
Disabled by default; set enabled: true to opt in. Prompt caching never changes
which model routing selects — it only keeps a conversation on the same
model/provider so the upstream prompt cache stays warm across turns, and
auto-injects provider cache-control markers where supported.
properties:
enabled:
type: boolean
description: "Master switch. Default false (opt-in). Applies across the entire instance."
session_affinity:
type: boolean
description: "Auto-derive a session key from the prompt prefix (system + tools + first user message) when X-Model-Affinity is absent, so follow-up turns reuse the same warm cache. Default true when enabled."
inject_cache_control:
type: boolean
description: "Auto-insert ephemeral cache_control breakpoints for providers that need explicit markers (Anthropic). Default true when enabled."
min_prefix_tokens:
type: integer
minimum: 1
description: "Skip breakpoint injection when the estimated stable prefix is below this many tokens. Default 1024."
session_ttl_seconds:
type: integer
minimum: 1
description: "Pin lifetime for implicit/explicit sessions; align with the provider cache window (e.g. 3600 for Anthropic 1h caching). Defaults to routing.session_ttl_seconds."
additionalProperties: false
system_prompt:
type: string
prompt_targets:
@ -510,6 +544,51 @@ properties:
Optional HTTP header name whose value is used as a tenant prefix in the cache key.
When set, keys are scoped as plano:affinity:{tenant_id}:{session_id}.
additionalProperties: false
routing_budget:
type: object
description: >
Per-session cost gate on model switching. Independent of prompt caching — it
applies whenever configured (presence of this block turns it on). The default
posture is to stick to the model a session is warm on. When routing proposes a
different model while that session's provider cache is plausibly still warm
(inferred from the idle gap vs. the provider's cache window), the actual
input-token cost of abandoning the cache —
context_tokens x (candidate_uncached_input_rate - anchor_cached_input_rate),
output cost deliberately excluded — accrues into the session's cumulative
switch spend. A paid switch is allowed only while that spend stays within
max_overhead_pct percent of the session's running never-switch baseline (what
staying on the anchor would have cost). An outright-cheaper switch is free but
never reduces the spend. Requires a cost source in model_metrics_sources.
properties:
max_overhead_pct:
type: number
minimum: 0
description: >
Cap on cumulative switching overhead, as a percentage of what the session
would have cost by never switching (a whole number: 20 = 20%). The promise
is "this conversation bills at most max_overhead_pct% above never-switching."
0 means never pay to switch (only outright-cheaper switches are allowed);
larger values buy more quality-driven switches. Typical range 10-30.
replenish_on_rebind:
type: boolean
description: "Reset the running baseline/spend totals when a cold session re-binds. Default true."
cache_read_discount:
type: number
minimum: 0
maximum: 1
description: >
Fallback used to estimate a model's cached input rate when the pricing feed
doesn't publish one (cached_rate = input_rate x discount). A pricing detail,
not a cost policy. Default 0.1.
record_counterfactual:
type: boolean
description: >
When true, a vetoed switch records the route the gate would have taken had
the switch been allowed, as the plano.switch.counterfactual_route span
attribute. Telemetry only — the counterfactual model is never dispatched.
Useful for evals/benchmarks. Default false.
required: ["max_overhead_pct"]
additionalProperties: false
additionalProperties: false
state_storage:
type: object