plano/Dockerfile

97 lines
4.1 KiB
Text
Raw Permalink Normal View History

2026-03-05 07:35:25 -08:00
# Envoy version — keep in sync with cli/planoai/consts.py ENVOY_VERSION
ARG ENVOY_VERSION=v1.37.0
# --- Dependency cache ---
FROM rust:1.93.0 AS deps
RUN rustup -v target add wasm32-wasip1
WORKDIR /arch
COPY crates/Cargo.toml crates/Cargo.lock ./
COPY crates/common/Cargo.toml common/Cargo.toml
COPY crates/hermesllm/Cargo.toml hermesllm/Cargo.toml
COPY crates/prompt_gateway/Cargo.toml prompt_gateway/Cargo.toml
COPY crates/llm_gateway/Cargo.toml llm_gateway/Cargo.toml
COPY crates/brightstaff/Cargo.toml brightstaff/Cargo.toml
# Dummy sources to pre-compile dependencies
RUN mkdir -p common/src && echo "" > common/src/lib.rs && \
mkdir -p hermesllm/src && echo "" > hermesllm/src/lib.rs && \
mkdir -p hermesllm/src/bin && echo "fn main() {}" > hermesllm/src/bin/fetch_models.rs && \
mkdir -p prompt_gateway/src && echo "#[no_mangle] pub fn _start() {}" > prompt_gateway/src/lib.rs && \
mkdir -p llm_gateway/src && echo "#[no_mangle] pub fn _start() {}" > llm_gateway/src/lib.rs && \
mkdir -p brightstaff/src && echo "fn main() {}" > brightstaff/src/main.rs && echo "" > brightstaff/src/lib.rs
RUN cargo build --release --target wasm32-wasip1 -p prompt_gateway -p llm_gateway || true
RUN cargo build --release -p brightstaff || true
# --- WASM plugins ---
FROM deps AS wasm-builder
RUN rm -rf common/src hermesllm/src prompt_gateway/src llm_gateway/src
COPY crates/common/src common/src
COPY crates/hermesllm/src hermesllm/src
COPY crates/prompt_gateway/src prompt_gateway/src
COPY crates/llm_gateway/src llm_gateway/src
RUN find common hermesllm prompt_gateway llm_gateway -name "*.rs" -exec touch {} +
RUN cargo build --release --target wasm32-wasip1 -p prompt_gateway -p llm_gateway
# --- Brightstaff binary ---
FROM deps AS brightstaff-builder
RUN rm -rf common/src hermesllm/src brightstaff/src
COPY crates/common/src common/src
COPY crates/hermesllm/src hermesllm/src
COPY crates/brightstaff/src brightstaff/src
RUN find common hermesllm brightstaff -name "*.rs" -exec touch {} +
RUN cargo build --release -p brightstaff
2026-03-05 07:35:25 -08:00
FROM docker.io/envoyproxy/envoy:${ENVOY_VERSION} AS envoy
FROM python:3.14-slim AS arch
feat(routing): automatic prompt caching + a per-session routing budget (#982) * feat(routing): cache-aware routing and zero-config prompt caching Reconcile prompt caching with intelligent routing: zero-config implicit session affinity, cache_control preservation/injection across transforms, cache-adjusted routing economics, and a response-driven cache-hit feedback loop. * refactor(caching): model-aware cache markers, drop cache-aware routing * feat(routing): session stickiness cache-regret cost gate * refactor(routing): expose per-model input and cached rates * fix(docker): patch libssh2 CVEs by upgrading after package install * feat: record counterfactual route on vetoed session switches Add opt-in session_stickiness.record_counterfactual that emits the plano.switch.counterfactual_route OTEL attribute when the cache-regret gate retains the previous model, capturing the route it would have taken had the switch been allowed. Telemetry only; the candidate is never dispatched. Includes schema, demo config, and a getting-started guide. * refactor: group prompt-caching and session-stickiness logic into cohesive modules Extract the two concerns interleaved in the LLM handler into dedicated sibling modules for readability: prompt_caching.rs (cache-marker injection) and session_stickiness.rs (session key/prefix-hash resolution, pin lookup, cache-regret gate, and pin planning). handlers/llm/mod.rs is now a thin orchestrator that calls them at clear phase boundaries. Behavior-preserving. * refactor(routing): unify session routing into a cumulative switch budget Replace observed_cache_hit + per-request threshold with time/provider-TTL warmth and a per-session switch budget in a single session_router::route() used by both the full-proxy and /routing decision paths. * refactor(routing): move switch budget to routing.routing_budget and rename telemetry * refactor(routing): drop negative-cost budget refund from routing_budget * refactor(routing): express routing_budget as a percentage overhead cap Replace the absolute seed_usd pool with max_overhead_pct (whole-number percent) measured against the session's running never-switch baseline: allow a paid switch only while cumulative switch spend stays within max_overhead_pct% of what staying on the anchor would have cost. Track baseline_usd and switch_spend_usd on the binding (monotonic, no refunds). * refactor(observability): align routing-budget telemetry with the overhead cap Rename the switch-decision reasons (within_budget/over_budget -> within_cap/ over_cap) and the session span attributes to match the percentage overhead-cap model: budget_remaining_in_usd -> overhead_pct + switch_spend_in_usd + baseline_in_usd, and switch threshold_in_usd -> overhead_ceiling_in_usd. * feat(observability): emit per-request and cumulative session cost Price each turn from the catalog rates and surface it as llm.usage.{input, output,total}_cost_usd on the (llm) span, then accumulate into a conversation- level session_cost_usd on the binding and emit plano.session.total_cost_in_usd on the routing span. Cache-creation tokens are priced at the plain input rate, and the OpenAI vs Anthropic prompt-token convention (cached folded in vs reported separately) is captured at parse time so the cost math is correct for both. Rates are resolved request-side since the response path is synchronous. * refactor(routing): price the never-switch baseline against the session default model Distinguish the session's default_model (the model it started on, i.e. what it would have cost by never switching) from anchor_model (the model that handled the latest request, which the session is warm on). The never-switch baseline now grows at the default model's cached rate rather than the drifting anchor's, so the overhead-cap denominator stays true to "% above never-switching"; switch cost is still measured against the current anchor. Also rename estimate_context_tokens -> actual_context_tokens (and est_context_tokens -> context_tokens) since the router prefers the provider's real prompt-token count when the session is warm. * feat(routing): track bounded route history and price returns to still-warm models Record a bounded (LRU, capped) per-model visit history on the session binding and use it to sharpen the switch-cost estimate: a return to a model still within its cache window re-reads only the tokens appended since its last visit (at its cached rate) instead of the whole context at the uncached rate, so an A->B->A switch is no longer over-charged as a full re-ingest. History is carried through the response-side refresh (refining the anchor's entry from real usage) and cleared on prefix drift. Adds the plano.switch.candidate_warm_tokens span attribute. * refactor(routing): tidy test names and rename pin events to binding events Shorten run-on test names and replace stale "pin" vocabulary in the session binding metric (brightstaff_session_pin_events_total -> _binding_events_total), dropping lifecycle event labels that are no longer emitted. * feat(routing): price switch cost against uncached rates when caching is off The overhead-cap gate previously always priced the anchor at its cached rate, assuming a warm provider cache. When prompt caching is disabled there is no warm cache to lose, so both the switch cost and the never-switch baseline are now priced at the plain uncached input rate: switch_cost = context_tokens x (candidate_uncached - anchor_uncached) / 1M. * fix(brightstaff): bound pricing-feed fetch with timeouts and warn when session state runs without a tenant header Co-authored-by: Cursor <cursoragent@cursor.com> * test(routing): validate the session overhead cap end-to-end across a multi-turn session Co-authored-by: Cursor <cursoragent@cursor.com> * docs(demos): fix stale pinning semantics and add routing_budget demo script Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Adil Hafeez <adil.hafeez@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-20 17:53:34 -07:00
# Install runtime deps first, then upgrade — so security patches also land on
# packages pulled in transitively by the install, not just those already present
# in the base image.
RUN set -eux; \
apt-get update; \
apt-get install -y --no-install-recommends gettext-base procps; \
feat(routing): automatic prompt caching + a per-session routing budget (#982) * feat(routing): cache-aware routing and zero-config prompt caching Reconcile prompt caching with intelligent routing: zero-config implicit session affinity, cache_control preservation/injection across transforms, cache-adjusted routing economics, and a response-driven cache-hit feedback loop. * refactor(caching): model-aware cache markers, drop cache-aware routing * feat(routing): session stickiness cache-regret cost gate * refactor(routing): expose per-model input and cached rates * fix(docker): patch libssh2 CVEs by upgrading after package install * feat: record counterfactual route on vetoed session switches Add opt-in session_stickiness.record_counterfactual that emits the plano.switch.counterfactual_route OTEL attribute when the cache-regret gate retains the previous model, capturing the route it would have taken had the switch been allowed. Telemetry only; the candidate is never dispatched. Includes schema, demo config, and a getting-started guide. * refactor: group prompt-caching and session-stickiness logic into cohesive modules Extract the two concerns interleaved in the LLM handler into dedicated sibling modules for readability: prompt_caching.rs (cache-marker injection) and session_stickiness.rs (session key/prefix-hash resolution, pin lookup, cache-regret gate, and pin planning). handlers/llm/mod.rs is now a thin orchestrator that calls them at clear phase boundaries. Behavior-preserving. * refactor(routing): unify session routing into a cumulative switch budget Replace observed_cache_hit + per-request threshold with time/provider-TTL warmth and a per-session switch budget in a single session_router::route() used by both the full-proxy and /routing decision paths. * refactor(routing): move switch budget to routing.routing_budget and rename telemetry * refactor(routing): drop negative-cost budget refund from routing_budget * refactor(routing): express routing_budget as a percentage overhead cap Replace the absolute seed_usd pool with max_overhead_pct (whole-number percent) measured against the session's running never-switch baseline: allow a paid switch only while cumulative switch spend stays within max_overhead_pct% of what staying on the anchor would have cost. Track baseline_usd and switch_spend_usd on the binding (monotonic, no refunds). * refactor(observability): align routing-budget telemetry with the overhead cap Rename the switch-decision reasons (within_budget/over_budget -> within_cap/ over_cap) and the session span attributes to match the percentage overhead-cap model: budget_remaining_in_usd -> overhead_pct + switch_spend_in_usd + baseline_in_usd, and switch threshold_in_usd -> overhead_ceiling_in_usd. * feat(observability): emit per-request and cumulative session cost Price each turn from the catalog rates and surface it as llm.usage.{input, output,total}_cost_usd on the (llm) span, then accumulate into a conversation- level session_cost_usd on the binding and emit plano.session.total_cost_in_usd on the routing span. Cache-creation tokens are priced at the plain input rate, and the OpenAI vs Anthropic prompt-token convention (cached folded in vs reported separately) is captured at parse time so the cost math is correct for both. Rates are resolved request-side since the response path is synchronous. * refactor(routing): price the never-switch baseline against the session default model Distinguish the session's default_model (the model it started on, i.e. what it would have cost by never switching) from anchor_model (the model that handled the latest request, which the session is warm on). The never-switch baseline now grows at the default model's cached rate rather than the drifting anchor's, so the overhead-cap denominator stays true to "% above never-switching"; switch cost is still measured against the current anchor. Also rename estimate_context_tokens -> actual_context_tokens (and est_context_tokens -> context_tokens) since the router prefers the provider's real prompt-token count when the session is warm. * feat(routing): track bounded route history and price returns to still-warm models Record a bounded (LRU, capped) per-model visit history on the session binding and use it to sharpen the switch-cost estimate: a return to a model still within its cache window re-reads only the tokens appended since its last visit (at its cached rate) instead of the whole context at the uncached rate, so an A->B->A switch is no longer over-charged as a full re-ingest. History is carried through the response-side refresh (refining the anchor's entry from real usage) and cleared on prefix drift. Adds the plano.switch.candidate_warm_tokens span attribute. * refactor(routing): tidy test names and rename pin events to binding events Shorten run-on test names and replace stale "pin" vocabulary in the session binding metric (brightstaff_session_pin_events_total -> _binding_events_total), dropping lifecycle event labels that are no longer emitted. * feat(routing): price switch cost against uncached rates when caching is off The overhead-cap gate previously always priced the anchor at its cached rate, assuming a warm provider cache. When prompt caching is disabled there is no warm cache to lose, so both the switch cost and the never-switch baseline are now priced at the plain uncached input rate: switch_cost = context_tokens x (candidate_uncached - anchor_uncached) / 1M. * fix(brightstaff): bound pricing-feed fetch with timeouts and warn when session state runs without a tenant header Co-authored-by: Cursor <cursoragent@cursor.com> * test(routing): validate the session overhead cap end-to-end across a multi-turn session Co-authored-by: Cursor <cursoragent@cursor.com> * docs(demos): fix stale pinning semantics and add routing_budget demo script Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Adil Hafeez <adil.hafeez@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-20 17:53:34 -07:00
apt-get upgrade -y; \
apt-get clean; rm -rf /var/lib/apt/lists/*
RUN pip install --no-cache-dir supervisor
# Remove PAM packages (CVE-2025-6020)
RUN set -eux; \
dpkg -r --force-depends libpam-modules libpam-modules-bin libpam-runtime libpam0g || true; \
dpkg -P --force-all libpam-modules libpam-modules-bin libpam-runtime libpam0g || true; \
rm -rf /etc/pam.d /lib/*/security /usr/lib/security || true
COPY --from=envoy /usr/local/bin/envoy /usr/local/bin/envoy
WORKDIR /app
2025-12-26 11:21:42 -08:00
RUN pip install --no-cache-dir uv
COPY cli/pyproject.toml ./
COPY cli/uv.lock ./
COPY cli/README.md ./
2026-03-05 07:35:25 -08:00
COPY config/plano_config_schema.yaml /config/plano_config_schema.yaml
COPY config/envoy.template.yaml /config/envoy.template.yaml
2025-12-26 11:21:42 -08:00
RUN pip install --no-cache-dir -e .
2025-12-26 11:21:42 -08:00
COPY cli/planoai planoai/
2025-12-25 14:55:29 -08:00
COPY config/envoy.template.yaml .
Rename all arch references to plano (#745) * Rename all arch references to plano across the codebase Complete rebrand from "Arch"/"archgw" to "Plano" including: - Config files: arch_config_schema.yaml, workflow, demo configs - Environment variables: ARCH_CONFIG_* → PLANO_CONFIG_* - Python CLI: variables, functions, file paths, docker mounts - Rust crates: config paths, log messages, metadata keys - Docker/build: Dockerfile, supervisord, .dockerignore, .gitignore - Docker Compose: volume mounts and env vars across all demos/tests - GitHub workflows: job/step names - Shell scripts: log messages - Demos: Python code, READMEs, VS Code configs, Grafana dashboard - Docs: RST includes, code comments, config references - Package metadata: package.json, pyproject.toml, uv.lock External URLs (docs.archgw.com, github.com/katanemo/archgw) left as-is. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Update remaining arch references in docs - Rename RST cross-reference labels: arch_access_logging, arch_overview_tracing, arch_overview_threading → plano_* - Update label references in request_lifecycle.rst - Rename arch_config_state_storage_example.yaml → plano_config_state_storage_example.yaml - Update config YAML comments: "Arch creates/uses" → "Plano creates/uses" - Update "the Arch gateway" → "the Plano gateway" in configuration_reference.rst - Update arch_config_schema.yaml reference in provider_models.py - Rename arch_agent_router → plano_agent_router in config example Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Fix remaining arch references found in second pass - config/docker-compose.dev.yaml: ARCH_CONFIG_FILE → PLANO_CONFIG_FILE, arch_config.yaml → plano_config.yaml, archgw_logs → plano_logs - config/test_passthrough.yaml: container mount path - tests/e2e/docker-compose.yaml: source file path (was still arch_config.yaml) - cli/planoai/core.py: comment and log message - crates/brightstaff/src/tracing/constants.rs: doc comment - tests/{e2e,archgw}/common.py: get_arch_messages → get_plano_messages, arch_state/arch_messages variables renamed - tests/{e2e,archgw}/test_prompt_gateway.py: updated imports and usages - demos/shared/test_runner/{common,test_demos}.py: same renames - tests/e2e/test_model_alias_routing.py: docstring - .dockerignore: archgw_modelserver → plano_modelserver - demos/use_cases/claude_code_router/pretty_model_resolution.sh: container name Note: x-arch-* HTTP header values and Rust constant names intentionally preserved for backwards compatibility with existing deployments. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-13 15:16:56 -08:00
COPY config/plano_config_schema.yaml .
RUN mkdir -p /etc/supervisor/conf.d
2025-12-25 14:55:29 -08:00
COPY config/supervisord.conf /etc/supervisor/conf.d/supervisord.conf
COPY --from=wasm-builder /arch/target/wasm32-wasip1/release/prompt_gateway.wasm /etc/envoy/proxy-wasm-plugins/prompt_gateway.wasm
COPY --from=wasm-builder /arch/target/wasm32-wasip1/release/llm_gateway.wasm /etc/envoy/proxy-wasm-plugins/llm_gateway.wasm
COPY --from=brightstaff-builder /arch/target/release/brightstaff /app/brightstaff
RUN mkdir -p /var/log/supervisor && \
touch /var/log/envoy.log /var/log/supervisor/supervisord.log \
/var/log/access_ingress.log /var/log/access_ingress_prompt.log \
/var/log/access_internal.log /var/log/access_llm.log /var/log/access_agent.log
ENTRYPOINT ["/usr/local/bin/supervisord", "-c", "/etc/supervisor/conf.d/supervisord.conf"]