plano/demos/getting_started/llm_gateway/config.yaml

version: v0.3.0

listeners:
  - type: model
    name: model_1
    address: 0.0.0.0
    port: 12000
    max_retries: 3

model_providers:

  - access_key: $OPENAI_API_KEY
    model: openai/gpt-4o-mini

  - access_key: $OPENAI_API_KEY
    model: openai/gpt-4.1

  - access_key: $OPENAI_API_KEY
    model: openai/gpt-4o
    default: true

  - access_key: $MISTRAL_API_KEY
    model: mistral/ministral-3b-latest

  - access_key: $ANTHROPIC_API_KEY
    model: anthropic/claude-3-7-sonnet-latest

  - access_key: $ANTHROPIC_API_KEY
    model: anthropic/claude-sonnet-4-0

  - access_key: $DEEPSEEK_API_KEY
    model: deepseek/deepseek-reasoner

  - access_key: $GROQ_API_KEY
    model: groq/llama-3.1-8b-instant

  - access_key: $GEMINI_API_KEY
    model: gemini/gemini-1.5-pro-latest

  - model: xai/grok-4-latest
    access_key: $GROK_API_KEY

  - model: together_ai/openai/gpt-oss-20b
    access_key: $TOGETHER_API_KEY

  - model: custom/test-model
    base_url: http://localhost:11223
    provider_interface: openai

tracing:
  random_sampling: 100
add support for agents (#564) 2025-10-14 14:01:11 -07:00			`version: v0.3.0`
Add support for streaming and fixes few issues (see description) (#202) 2024-10-28 20:05:06 -04:00
Update arch_config and add tests for arch config file (#407) 2025-02-14 19:28:10 -08:00			`listeners:`
add support for agents (#564) 2025-10-14 14:01:11 -07:00			`- type: model`
			`name: model_1`
Update arch_config and add tests for arch config file (#407) 2025-02-14 19:28:10 -08:00			`address: 0.0.0.0`
			`port: 12000`
add envoy retries (#712) * add envoy retries * add missing file * fix tests --------- Co-authored-by: Adil Hafeez <adil.hafeez10@t-mobile.com> 2026-01-28 20:31:01 -08:00			`max_retries: 3`
Add support for streaming and fixes few issues (see description) (#202) 2024-10-28 20:05:06 -04:00
add support for agents (#564) 2025-10-14 14:01:11 -07:00			`model_providers:`
add support for openwebui (#487) 2025-05-28 19:08:00 -07:00
better model names (#517) 2025-07-11 16:42:16 -07:00			`- access_key: $OPENAI_API_KEY`
			`model: openai/gpt-4o-mini`
Add support for streaming and fixes few issues (see description) (#202) 2024-10-28 20:05:06 -04:00
bug fix - allow image content to pass through (#539) fixes https://github.com/katanemo/archgw/issues/535 2025-07-25 01:22:06 -07:00			`- access_key: $OPENAI_API_KEY`
			`model: openai/gpt-4.1`

better model names (#517) 2025-07-11 16:42:16 -07:00			`- access_key: $OPENAI_API_KEY`
			`model: openai/gpt-4o`
add support for openwebui (#487) 2025-05-28 19:08:00 -07:00			`default: true`
Add support for streaming and fixes few issues (see description) (#202) 2024-10-28 20:05:06 -04:00
better model names (#517) 2025-07-11 16:42:16 -07:00			`- access_key: $MISTRAL_API_KEY`
			`model: mistral/ministral-3b-latest`

			`- access_key: $ANTHROPIC_API_KEY`
add support for v1/messages and transformations (#558) * pushing draft PR * transformations are working. Now need to add some tests next * updated tests and added necessary response transformations for Anthropics' message response object * fixed bugs for integration tests * fixed doc tests * fixed serialization issues with enums on response * adding some debug logs to help * fixed issues with non-streaming responses * updated the stream_context to update response bytes * the serialized bytes length must be set in the response side * fixed the debug statement that was causing the integration tests for wasm to fail * fixing json parsing errors * intentionally removing the headers * making sure that we convert the raw bytes to the correct provider type upstream * fixing non-streaming responses to tranform correctly * /v1/messages works with transformations to and from /v1/chat/completions * updating the CLI and demos to support anthropic vs. claude * adding the anthropic key to the preference based routing tests * fixed test cases and added more structured logs * fixed integration tests and cleaned up logs * added python client tests for anthropic and openai * cleaned up logs and fixed issue with connectivity for llm gateway in weather forecast demo * fixing the tests. python dependency order was broken * updated the openAI client to fix demos * removed the raw response debug statement * fixed the dup cloning issue and cleaned up the ProviderRequestType enum and traits * fixing logs * moved away from string literals to consts * fixed streaming from Anthropic Client to OpenAI * removed debug statement that would likely trip up integration tests * fixed integration tests for llm_gateway * cleaned up test cases and removed unnecessary crates * fixing comments from PR * fixed bug whereby we were sending an OpenAIChatCompletions request object to llm_gateway even though the request may have been AnthropicMessages --------- Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-4.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-9.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-10.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-41.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-136.local> 2025-09-10 07:40:30 -07:00			`model: anthropic/claude-3-7-sonnet-latest`
better model names (#517) 2025-07-11 16:42:16 -07:00
			`- access_key: $ANTHROPIC_API_KEY`
add support for v1/messages and transformations (#558) * pushing draft PR * transformations are working. Now need to add some tests next * updated tests and added necessary response transformations for Anthropics' message response object * fixed bugs for integration tests * fixed doc tests * fixed serialization issues with enums on response * adding some debug logs to help * fixed issues with non-streaming responses * updated the stream_context to update response bytes * the serialized bytes length must be set in the response side * fixed the debug statement that was causing the integration tests for wasm to fail * fixing json parsing errors * intentionally removing the headers * making sure that we convert the raw bytes to the correct provider type upstream * fixing non-streaming responses to tranform correctly * /v1/messages works with transformations to and from /v1/chat/completions * updating the CLI and demos to support anthropic vs. claude * adding the anthropic key to the preference based routing tests * fixed test cases and added more structured logs * fixed integration tests and cleaned up logs * added python client tests for anthropic and openai * cleaned up logs and fixed issue with connectivity for llm gateway in weather forecast demo * fixing the tests. python dependency order was broken * updated the openAI client to fix demos * removed the raw response debug statement * fixed the dup cloning issue and cleaned up the ProviderRequestType enum and traits * fixing logs * moved away from string literals to consts * fixed streaming from Anthropic Client to OpenAI * removed debug statement that would likely trip up integration tests * fixed integration tests for llm_gateway * cleaned up test cases and removed unnecessary crates * fixing comments from PR * fixed bug whereby we were sending an OpenAIChatCompletions request object to llm_gateway even though the request may have been AnthropicMessages --------- Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-4.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-9.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-10.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-41.local> Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-136.local> 2025-09-10 07:40:30 -07:00			`model: anthropic/claude-sonnet-4-0`
better model names (#517) 2025-07-11 16:42:16 -07:00
			`- access_key: $DEEPSEEK_API_KEY`
			`model: deepseek/deepseek-reasoner`

			`- access_key: $GROQ_API_KEY`
			`model: groq/llama-3.1-8b-instant`

			`- access_key: $GEMINI_API_KEY`
			`model: gemini/gemini-1.5-pro-latest`

draft commit to add support for xAI, TogehterAI, AzureOpenAI (#570) * draft commit to add support for xAI, LambdaAI, TogehterAI, AzureOpenAI * fixing failing tests and updating rederend config file * Update arch_config_with_aliases.yaml * adding the AZURE_API_KEY to the GH workflow for e2e * fixing GH secerts * adding valdiating for azure_openai --------- Co-authored-by: Salman Paracha <salmanparacha@MacBook-Pro-167.local> 2025-09-18 18:36:30 -07:00			`- model: xai/grok-4-latest`
			`access_key: $GROK_API_KEY`

			`- model: together_ai/openai/gpt-oss-20b`
			`access_key: $TOGETHER_API_KEY`

better model names (#517) 2025-07-11 16:42:16 -07:00			`- model: custom/test-model`
Run plano natively by default (#744) 2026-03-05 07:35:25 -08:00			`base_url: http://localhost:11223`
better model names (#517) 2025-07-11 16:42:16 -07:00			`provider_interface: openai`
add support for gemini (#505) 2025-06-11 15:15:00 -07:00
Add support for streaming and fixes few issues (see description) (#202) 2024-10-28 20:05:06 -04:00			`tracing:`
			`random_sampling: 100`