mirror of
https://github.com/katanemo/plano.git
synced 2026-07-23 16:51:04 +02:00
* feat(routing): cache-aware routing and zero-config prompt caching
Reconcile prompt caching with intelligent routing: zero-config implicit
session affinity, cache_control preservation/injection across transforms,
cache-adjusted routing economics, and a response-driven cache-hit feedback loop.
* refactor(caching): model-aware cache markers, drop cache-aware routing
* feat(routing): session stickiness cache-regret cost gate
* refactor(routing): expose per-model input and cached rates
* fix(docker): patch libssh2 CVEs by upgrading after package install
* feat: record counterfactual route on vetoed session switches
Add opt-in session_stickiness.record_counterfactual that emits the
plano.switch.counterfactual_route OTEL attribute when the cache-regret
gate retains the previous model, capturing the route it would have taken
had the switch been allowed. Telemetry only; the candidate is never
dispatched. Includes schema, demo config, and a getting-started guide.
* refactor: group prompt-caching and session-stickiness logic into cohesive modules
Extract the two concerns interleaved in the LLM handler into dedicated
sibling modules for readability: prompt_caching.rs (cache-marker injection)
and session_stickiness.rs (session key/prefix-hash resolution, pin lookup,
cache-regret gate, and pin planning). handlers/llm/mod.rs is now a thin
orchestrator that calls them at clear phase boundaries. Behavior-preserving.
* refactor(routing): unify session routing into a cumulative switch budget
Replace observed_cache_hit + per-request threshold with time/provider-TTL
warmth and a per-session switch budget in a single session_router::route()
used by both the full-proxy and /routing decision paths.
* refactor(routing): move switch budget to routing.routing_budget and rename telemetry
* refactor(routing): drop negative-cost budget refund from routing_budget
* refactor(routing): express routing_budget as a percentage overhead cap
Replace the absolute seed_usd pool with max_overhead_pct (whole-number
percent) measured against the session's running never-switch baseline:
allow a paid switch only while cumulative switch spend stays within
max_overhead_pct% of what staying on the anchor would have cost. Track
baseline_usd and switch_spend_usd on the binding (monotonic, no refunds).
* refactor(observability): align routing-budget telemetry with the overhead cap
Rename the switch-decision reasons (within_budget/over_budget -> within_cap/
over_cap) and the session span attributes to match the percentage overhead-cap
model: budget_remaining_in_usd -> overhead_pct + switch_spend_in_usd +
baseline_in_usd, and switch threshold_in_usd -> overhead_ceiling_in_usd.
* feat(observability): emit per-request and cumulative session cost
Price each turn from the catalog rates and surface it as llm.usage.{input,
output,total}_cost_usd on the (llm) span, then accumulate into a conversation-
level session_cost_usd on the binding and emit plano.session.total_cost_in_usd
on the routing span. Cache-creation tokens are priced at the plain input rate,
and the OpenAI vs Anthropic prompt-token convention (cached folded in vs
reported separately) is captured at parse time so the cost math is correct for
both. Rates are resolved request-side since the response path is synchronous.
* refactor(routing): price the never-switch baseline against the session default model
Distinguish the session's default_model (the model it started on, i.e. what it
would have cost by never switching) from anchor_model (the model that handled
the latest request, which the session is warm on). The never-switch baseline now
grows at the default model's cached rate rather than the drifting anchor's, so
the overhead-cap denominator stays true to "% above never-switching"; switch cost
is still measured against the current anchor. Also rename estimate_context_tokens
-> actual_context_tokens (and est_context_tokens -> context_tokens) since the
router prefers the provider's real prompt-token count when the session is warm.
* feat(routing): track bounded route history and price returns to still-warm models
Record a bounded (LRU, capped) per-model visit history on the session binding
and use it to sharpen the switch-cost estimate: a return to a model still within
its cache window re-reads only the tokens appended since its last visit (at its
cached rate) instead of the whole context at the uncached rate, so an A->B->A
switch is no longer over-charged as a full re-ingest. History is carried through
the response-side refresh (refining the anchor's entry from real usage) and
cleared on prefix drift. Adds the plano.switch.candidate_warm_tokens span
attribute.
* refactor(routing): tidy test names and rename pin events to binding events
Shorten run-on test names and replace stale "pin" vocabulary in the session
binding metric (brightstaff_session_pin_events_total -> _binding_events_total),
dropping lifecycle event labels that are no longer emitted.
* feat(routing): price switch cost against uncached rates when caching is off
The overhead-cap gate previously always priced the anchor at its cached rate,
assuming a warm provider cache. When prompt caching is disabled there is no
warm cache to lose, so both the switch cost and the never-switch baseline are
now priced at the plain uncached input rate:
switch_cost = context_tokens x (candidate_uncached - anchor_uncached) / 1M.
* fix(brightstaff): bound pricing-feed fetch with timeouts and warn when session state runs without a tenant header
Co-authored-by: Cursor <cursoragent@cursor.com>
* test(routing): validate the session overhead cap end-to-end across a multi-turn session
Co-authored-by: Cursor <cursoragent@cursor.com>
* docs(demos): fix stale pinning semantics and add routing_budget demo script
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: Adil Hafeez <adil.hafeez@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
96 lines
4.1 KiB
Docker
96 lines
4.1 KiB
Docker
# Envoy version — keep in sync with cli/planoai/consts.py ENVOY_VERSION
|
|
ARG ENVOY_VERSION=v1.37.0
|
|
|
|
# --- Dependency cache ---
|
|
FROM rust:1.93.0 AS deps
|
|
RUN rustup -v target add wasm32-wasip1
|
|
WORKDIR /arch
|
|
|
|
COPY crates/Cargo.toml crates/Cargo.lock ./
|
|
COPY crates/common/Cargo.toml common/Cargo.toml
|
|
COPY crates/hermesllm/Cargo.toml hermesllm/Cargo.toml
|
|
COPY crates/prompt_gateway/Cargo.toml prompt_gateway/Cargo.toml
|
|
COPY crates/llm_gateway/Cargo.toml llm_gateway/Cargo.toml
|
|
COPY crates/brightstaff/Cargo.toml brightstaff/Cargo.toml
|
|
|
|
# Dummy sources to pre-compile dependencies
|
|
RUN mkdir -p common/src && echo "" > common/src/lib.rs && \
|
|
mkdir -p hermesllm/src && echo "" > hermesllm/src/lib.rs && \
|
|
mkdir -p hermesllm/src/bin && echo "fn main() {}" > hermesllm/src/bin/fetch_models.rs && \
|
|
mkdir -p prompt_gateway/src && echo "#[no_mangle] pub fn _start() {}" > prompt_gateway/src/lib.rs && \
|
|
mkdir -p llm_gateway/src && echo "#[no_mangle] pub fn _start() {}" > llm_gateway/src/lib.rs && \
|
|
mkdir -p brightstaff/src && echo "fn main() {}" > brightstaff/src/main.rs && echo "" > brightstaff/src/lib.rs
|
|
|
|
RUN cargo build --release --target wasm32-wasip1 -p prompt_gateway -p llm_gateway || true
|
|
RUN cargo build --release -p brightstaff || true
|
|
|
|
# --- WASM plugins ---
|
|
FROM deps AS wasm-builder
|
|
RUN rm -rf common/src hermesllm/src prompt_gateway/src llm_gateway/src
|
|
COPY crates/common/src common/src
|
|
COPY crates/hermesllm/src hermesllm/src
|
|
COPY crates/prompt_gateway/src prompt_gateway/src
|
|
COPY crates/llm_gateway/src llm_gateway/src
|
|
RUN find common hermesllm prompt_gateway llm_gateway -name "*.rs" -exec touch {} +
|
|
RUN cargo build --release --target wasm32-wasip1 -p prompt_gateway -p llm_gateway
|
|
|
|
# --- Brightstaff binary ---
|
|
FROM deps AS brightstaff-builder
|
|
RUN rm -rf common/src hermesllm/src brightstaff/src
|
|
COPY crates/common/src common/src
|
|
COPY crates/hermesllm/src hermesllm/src
|
|
COPY crates/brightstaff/src brightstaff/src
|
|
RUN find common hermesllm brightstaff -name "*.rs" -exec touch {} +
|
|
RUN cargo build --release -p brightstaff
|
|
|
|
FROM docker.io/envoyproxy/envoy:${ENVOY_VERSION} AS envoy
|
|
|
|
FROM python:3.14-slim AS arch
|
|
|
|
# Install runtime deps first, then upgrade — so security patches also land on
|
|
# packages pulled in transitively by the install, not just those already present
|
|
# in the base image.
|
|
RUN set -eux; \
|
|
apt-get update; \
|
|
apt-get install -y --no-install-recommends gettext-base procps; \
|
|
apt-get upgrade -y; \
|
|
apt-get clean; rm -rf /var/lib/apt/lists/*
|
|
|
|
RUN pip install --no-cache-dir supervisor
|
|
|
|
# Remove PAM packages (CVE-2025-6020)
|
|
RUN set -eux; \
|
|
dpkg -r --force-depends libpam-modules libpam-modules-bin libpam-runtime libpam0g || true; \
|
|
dpkg -P --force-all libpam-modules libpam-modules-bin libpam-runtime libpam0g || true; \
|
|
rm -rf /etc/pam.d /lib/*/security /usr/lib/security || true
|
|
|
|
COPY --from=envoy /usr/local/bin/envoy /usr/local/bin/envoy
|
|
|
|
WORKDIR /app
|
|
|
|
RUN pip install --no-cache-dir uv
|
|
|
|
COPY cli/pyproject.toml ./
|
|
COPY cli/uv.lock ./
|
|
COPY cli/README.md ./
|
|
COPY config/plano_config_schema.yaml /config/plano_config_schema.yaml
|
|
COPY config/envoy.template.yaml /config/envoy.template.yaml
|
|
|
|
RUN pip install --no-cache-dir -e .
|
|
|
|
COPY cli/planoai planoai/
|
|
COPY config/envoy.template.yaml .
|
|
COPY config/plano_config_schema.yaml .
|
|
RUN mkdir -p /etc/supervisor/conf.d
|
|
COPY config/supervisord.conf /etc/supervisor/conf.d/supervisord.conf
|
|
|
|
COPY --from=wasm-builder /arch/target/wasm32-wasip1/release/prompt_gateway.wasm /etc/envoy/proxy-wasm-plugins/prompt_gateway.wasm
|
|
COPY --from=wasm-builder /arch/target/wasm32-wasip1/release/llm_gateway.wasm /etc/envoy/proxy-wasm-plugins/llm_gateway.wasm
|
|
COPY --from=brightstaff-builder /arch/target/release/brightstaff /app/brightstaff
|
|
|
|
RUN mkdir -p /var/log/supervisor && \
|
|
touch /var/log/envoy.log /var/log/supervisor/supervisord.log \
|
|
/var/log/access_ingress.log /var/log/access_ingress_prompt.log \
|
|
/var/log/access_internal.log /var/log/access_llm.log /var/log/access_agent.log
|
|
|
|
ENTRYPOINT ["/usr/local/bin/supervisord", "-c", "/etc/supervisor/conf.d/supervisord.conf"]
|