Merge branch 'main' into adil/add_acm_demo

2026-06-17 15:25:17 +02:00 · 2024-12-06 16:17:26 -08:00 · 2024-12-06 16:17:26 -08:00 · 285a66fdb6
commit 285a66fdb6
parent 5e182b6c09 af0e7d178b
40 changed files with 5270 additions and 1124 deletions
--- a/README.md
+++ b/README.md
@ -33,94 +33,184 @@ Engineered with purpose-built LLMs, Arch handles the critical but undifferentiat
 To get in touch with us, please join our [discord server](https://discord.gg/pGZf2gcwEc). We will be monitoring that actively and offering support there.

 ## Demos
-* [Weather Forecast](demos/weather_forecast/README.md) - Walk through of the core function calling capabilities of of arch gateway using weather forecasting service
+* [Weather Forecast](demos/weather_forecast/README.md) - Walk through of the core function calling capabilities of arch gateway using weather forecasting service
 * [Insurance Agent](demos/insurance_agent/README.md) - Build a full insurance agent with Arch
 * [Network Agent](demos/network_agent/README.md) - Build a networking co-pilot/agent agent with Arch

 ## Quickstart

-Follow this guide to learn how to quickly set up Arch and integrate it into your generative AI applications.
+Follow this quickstart guide to use arch gateway to build a simple AI agent. Laster in the section we will see how you can Arch Gateway to manage access keys, provide unified access to upstream LLMs and to provide e2e observability.

 ### Prerequisites

 Before you begin, ensure you have the following:

- `Docker` & `Python` installed on your system
- `API Keys` for LLM providers (if using external LLMs)
-
-
-### Step 1: Install Arch
-
-Arch's CLI allows you to manage and interact with the Arch gateway efficiently. To install the CLI, simply run the following command:
-Tip: We recommend that developers create a new Python virtual environment to isolate dependencies before installing Arch. This ensures that archgw and its dependencies do not interfere with other packages on your system.
-
-Make sure you have following utilities installed before proceeding further,
-
 1. [Docker System](https://docs.docker.com/get-started/get-docker/) (v24)
 2. [Docker compose](https://docs.docker.com/compose/install/) (v2.29)
 3. [Python](https://www.python.org/downloads/) (v3.12)
-4. [Poetry](https://python-poetry.org/docs/#installing-with-the-official-installer) (v1.8.3. *Note: only needed for local development*)

+Arch's CLI allows you to manage and interact with the Arch gateway efficiently. To install the CLI, simply run the following command:
+
+> [!TIP]
+> We recommend that developers create a new Python virtual environment to isolate dependencies before installing Arch. This ensures that archgw and its dependencies do not interfere with other packages on your system.

 ```console
 $ python -m venv venv
 $ source venv/bin/activate   # On Windows, use: venv\Scripts\activate
-$ pip install archgw
+$ pip install archgw==0.1.6
 ```

-### Step 2: Configure Arch with your application
+### Build AI Agent with Arch Gateway

-Arch operates based on a configuration file where you can define LLM providers, prompt targets, guardrails, etc.
-Below is an example configuration to get you started:
+In following quickstart we will show you how easy it is to build AI agent with Arch gateway. We will build a currency exchange agent using following simple steps. For this demo we will use `https://api.frankfurter.dev/` to fetch latest price for currencies and assume USD as base currency.
+
+#### Step 1. Create arch config file
+
+Create `arch_config.yaml` file with following content,

 ```yaml
 version: v0.1
+
 listener:
-  address: 127.0.0.1
-  port: 8080 #If you configure port 443, you'll need to update the listener with tls_certificates
+  address: 0.0.0.0
+  port: 10000
  message_format: huggingface
+  connect_timeout: 0.005s

-# Centralized way to manage LLMs, manage keys, retry logic, failover and limits in a central way
 llm_providers:
-  - name: OpenAI
-    provider: openai
+  - name: gpt-4o
    access_key: $OPENAI_API_KEY
-    model: gpt-3.5-turbo
-    default: true
+    provider: openai
+    model: gpt-4o

-# default system prompt used by all prompt targets
 system_prompt: |
-  You are a network assistant that helps operators with a better understanding of network traffic flow and perform actions on networking operations. No advice on manufacturers or purchasing decisions.
+  You are a helpful assistant.
+
+prompt_guards:
+  input_guards:
+    jailbreak:
+      on_exception:
+        message: Looks like you're curious about my abilities, but I can only provide assistance for currency exchange.

 prompt_targets:
-    - name: device_summary
-      description: Retrieve network statistics for specific devices within a time range
-      endpoint:
-        name: app_server
-        path: /agent/device_summary
-      parameters:
-        - name: device_ids
-          type: list
-          description: A list of device identifiers (IDs) to retrieve statistics for.
-          required: true  # device_ids are required to get device statistics
-        - name: days
-          type: int
-          description: The number of days for which to gather device statistics.
-          default: "7"
+  - name: currency_exchange
+    description: Get currency exchange rate from USD to other currencies
+    parameters:
+      - name: currency_symbol
+        description: the currency that needs conversion
+        required: true
+        type: str
+        in_path: true
+    endpoint:
+      name: frankfurther_api
+      path: /v1/latest?base=USD&symbols={currency_symbol}
+    system_prompt: |
+      You are a helpful assistant. Show me the currency symbol you want to convert from USD.
+
+  - name: get_supported_currencies
+    description: Get list of supported currencies for conversion
+    endpoint:
+      name: frankfurther_api
+      path: /v1/currencies

-# Arch creates a round-robin load balancing between different endpoints, managed via the cluster subsystem.
 endpoints:
-  app_server:
-    # value could be ip address or a hostname with port
-    # this could also be a list of endpoints for load balancing
-    # for example endpoint: [ ip1:port, ip2:port ]
-    endpoint: host.docker.internal:18083
-    # max time to wait for a connection to be established
-    connect_timeout: 0.005s
+  frankfurther_api:
+    endpoint: api.frankfurter.dev:443
+    protocol: https
 ```
-### Step 3: Using OpenAI Client with Arch as an Egress Gateway

-Make outbound calls via Arch
+#### Step 2. Start arch gateway with currency conversion config
+
+```sh
+
+$ archgw up arch_config.yaml
+2024-12-05 16:56:27,979 - cli.main - INFO - Starting archgw cli version: 0.1.5
+...
+2024-12-05 16:56:28,485 - cli.utils - INFO - Schema validation successful!
+2024-12-05 16:56:28,485 - cli.main - INFO - Starging arch model server and arch gateway
+...
+2024-12-05 16:56:51,647 - cli.core - INFO - Container is healthy!
+
+```
+
+Once the gateway is up you can start interacting with at port 10000 using openai chat completion API.
+
+Some of the sample queries you can ask could be `what is currency rate for gbp?` or `show me list of currencies for conversion`.
+
+#### Step 3. Interacting with gateway using curl command
+
+Here is a sample curl command you can use to interact,
+
+```bash
+$ curl --header 'Content-Type: application/json' \
+  --data '{"messages": [{"role": "user","content": "what is exchange rate for gbp"}]}' \
+  http://localhost:10000/v1/chat/completions | jq ".choices[0].message.content"
+
+"As of the date provided in your context, December 5, 2024, the exchange rate for GBP (British Pound) from USD (United States Dollar) is 0.78558. This means that 1 USD is equivalent to 0.78558 GBP."
+
+```
+
+And to get list of supported currencies,
+
+```bash
+$ curl --header 'Content-Type: application/json' \
+  --data '{"messages": [{"role": "user","content": "show me list of currencies that are supported for conversion"}]}' \
+  http://localhost:10000/v1/chat/completions | jq ".choices[0].message.content"
+
+"Here is a list of the currencies that are supported for conversion from USD, along with their symbols:\n\n1. AUD - Australian Dollar\n2. BGN - Bulgarian Lev\n3. BRL - Brazilian Real\n4. CAD - Canadian Dollar\n5. CHF - Swiss Franc\n6. CNY - Chinese Renminbi Yuan\n7. CZK - Czech Koruna\n8. DKK - Danish Krone\n9. EUR - Euro\n10. GBP - British Pound\n11. HKD - Hong Kong Dollar\n12. HUF - Hungarian Forint\n13. IDR - Indonesian Rupiah\n14. ILS - Israeli New Sheqel\n15. INR - Indian Rupee\n16. ISK - Icelandic Króna\n17. JPY - Japanese Yen\n18. KRW - South Korean Won\n19. MXN - Mexican Peso\n20. MYR - Malaysian Ringgit\n21. NOK - Norwegian Krone\n22. NZD - New Zealand Dollar\n23. PHP - Philippine Peso\n24. PLN - Polish Złoty\n25. RON - Romanian Leu\n26. SEK - Swedish Krona\n27. SGD - Singapore Dollar\n28. THB - Thai Baht\n29. TRY - Turkish Lira\n30. USD - United States Dollar\n31. ZAR - South African Rand\n\nIf you want to convert USD to any of these currencies, you can select the one you are interested in."
+
+```
+
+### Use Arch Gateway as LLM Router
+
+#### Step 1. Create arch config file
+
+Arch operates based on a configuration file where you can define LLM providers, prompt targets, guardrails, etc. Below is an example configuration that defines openai and mistral LLM providers.
+
+Create `arch_config.yaml` file with following content:
+
+```yaml
+version: v0.1
+
+listener:
+  address: 0.0.0.0
+  port: 10000
+  message_format: huggingface
+  connect_timeout: 0.005s
+
+llm_providers:
+  - name: gpt-4o
+    access_key: $OPENAI_API_KEY
+    provider: openai
+    model: gpt-4o
+    default: true
+
+  - name: ministral-3b
+    access_key: $MISTRAL_API_KEY
+    provider: mistral
+    model: ministral-3b-latest
+```
+
+#### Step 2. Start arch gateway
+
+Once the config file is created ensure that you have env vars setup for `MISTRAL_API_KEY` and `OPENAI_API_KEY` (or these are defined in `.env` file).
+
+Start arch gateway,
+
+```
+$ archgw up arch_config.yaml
+2024-12-05 11:24:51,288 - cli.main - INFO - Starting archgw cli version: 0.1.5
+2024-12-05 11:24:51,825 - cli.utils - INFO - Schema validation successful!
+2024-12-05 11:24:51,825 - cli.main - INFO - Starting arch model server and arch gateway
+...
+2024-12-05 11:25:16,131 - cli.core - INFO - Container is healthy!
+```
+
+### Step 3: Interact with LLM
+
+#### Step 3.1: Using OpenAI python client
+
+Make outbound calls via Arch gateway

 ```python
 from openai import OpenAI
@ -143,12 +233,59 @@ print("OpenAI Response:", response.choices[0].message.content)

 ```

-### [Observability](https://docs.archgw.com/guides/observability/observability.html)
+#### Step 3.2: Using curl command
+```
+$ curl --header 'Content-Type: application/json' \
+  --data '{"messages": [{"role": "user","content": "What is the capital of France?"}]}' \
+  http://localhost:12000/v1/chat/completions
+
+{
+  ...
+  "model": "gpt-4o-2024-08-06",
+  "choices": [
+    {
+      ...
+      "message": {
+        "role": "assistant",
+        "content": "The capital of France is Paris.",
+      },
+    }
+  ],
+...
+}
+
+```
+
+You can override model selection using `x-arch-llm-provider-hint` header. For example if you want to use mistral using following curl command,
+
+```
+$ curl --header 'Content-Type: application/json' \
+  --header 'x-arch-llm-provider-hint: ministral-3b' \
+  --data '{"messages": [{"role": "user","content": "What is the capital of France?"}]}' \
+  http://localhost:12000/v1/chat/completions
+{
+  ...
+  "model": "ministral-3b-latest",
+  "choices": [
+    {
+      "message": {
+        "role": "assistant",
+        "content": "The capital of France is Paris. It is the most populous city in France and is known for its iconic landmarks such as the Eiffel Tower, the Louvre Museum, and Notre-Dame Cathedral. Paris is also a major global center for art, fashion, gastronomy, and culture.",
+      },
+      ...
+    }
+  ],
+  ...
+}
+
+```
+
+## [Observability](https://docs.archgw.com/guides/observability/observability.html)
 Arch is designed to support best-in class observability by supporting open standards. Please read our [docs](https://docs.archgw.com/guides/observability/observability.html) on observability for more details on tracing, metrics, and logs. The screenshot below is from our integration with Signoz (among others)

 ![alt text](docs/source/_static/img/tracing.png)

-### Contribution
+## Contribution
 We would love feedback on our [Roadmap](https://github.com/orgs/katanemo/projects/1) and we welcome contributions to **Arch**!
 Whether you're fixing bugs, adding new features, improving documentation, or creating tutorials, your help is much appreciated.
 Please visit our [Contribution Guide](CONTRIBUTING.md) for more details
--- a/arch/arch_config_schema.yaml
+++ b/arch/arch_config_schema.yaml
@ -28,6 +28,11 @@ properties:
            type: string
          connect_timeout:
            type: string
+          protocol:
+            type: string
+            enum:
+              - http
+              - https
        additionalProperties: false
        required:
          - endpoint
@ -106,6 +111,11 @@ properties:
              type: string
            path:
              type: string
+            http_method:
+              type: string
+              enum:
+                - GET
+                - POST
          additionalProperties: false
          required:
            - name
--- a/arch/envoy.template.yaml
+++ b/arch/envoy.template.yaml
@ -500,7 +500,18 @@ static_resources:
                    socket_address:
                      address: {{ cluster.endpoint }}
                      port_value: {{ cluster.port }}
-                  hostname: {{ cluster.name }}
+                  hostname: {{ cluster.endpoint }}
+      {% if cluster.protocol == "https" %}
+      transport_socket:
+        name: envoy.transport_sockets.tls
+        typed_config:
+          "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.UpstreamTlsContext
+          sni: {{ cluster.endpoint }}
+          common_tls_context:
+            tls_params:
+              tls_minimum_protocol_version: TLSv1_2
+              tls_maximum_protocol_version: TLSv1_3
+      {% endif %}
 {% endfor %}
    - name: arch_internal
      connect_timeout: 5s
--- a/arch/tools/README.md
+++ b/arch/tools/README.md
@ -19,7 +19,7 @@ source venv/bin/activate

 ### Step 3: Run the build script
 ```bash
-pip install archgw
+pip install archgw==0.1.6
 ```

 ## Uninstall Instructions: archgw CLI
@ -48,7 +48,7 @@ source venv/bin/activate

 ### Step 3: Run the build script
 ```bash
-sh build_cli.sh
+poetry install
 ```

 ### Step 4: build Arch
--- a/arch/tools/cli/core.py
+++ b/arch/tools/cli/core.py
@ -196,7 +196,7 @@ def start_arch_modelserver():
        subprocess.run(
            ["archgw_modelserver", "restart"], check=True, start_new_session=True
        )
-        log.info("Successfull ran model_server")
+        log.info("Successfully ran model_server")
    except subprocess.CalledProcessError as e:
        log.info(f"Failed to start model_server. Please check archgw_modelserver logs")
        sys.exit(1)
@ -212,7 +212,7 @@ def stop_arch_modelserver():
            ["archgw_modelserver", "stop"],
            check=True,
        )
-        log.info("Successfull stopped the archgw model_server")
+        log.info("Successfully stopped the archgw model_server")
    except subprocess.CalledProcessError as e:
        log.info(f"Failed to start model_server. Please check archgw_modelserver logs")
        sys.exit(1)
--- a/arch/tools/cli/main.py
+++ b/arch/tools/cli/main.py
@ -171,7 +171,7 @@ def up(file, path, service):
        log.info(f"Error: {str(e)}")
        sys.exit(1)

-    log.info("Starging arch model server and arch gateway")
+    log.info("Starting arch model server and arch gateway")

    # Set the ARCH_CONFIG_FILE environment variable
    env_stage = {}
--- a/arch/tools/poetry.lock
+++ b/arch/tools/poetry.lock
--- a/arch/tools/pyproject.toml
+++ b/arch/tools/pyproject.toml
@ -1,6 +1,6 @@
 [tool.poetry]
 name = "archgw"
-version = "0.1.5"
+version = "0.1.6"
 description = "Python-based CLI tool to manage Arch Gateway."
 authors = ["Katanemo Labs, Inc."]
 packages = [
@ -10,7 +10,7 @@ readme = "README.md"

 [tool.poetry.dependencies]
 python = ">=3.12"
-archgw_modelserver = "0.1.5"
+archgw_modelserver = "0.1.6"
 pyyaml = "^6.0.2"
 pydantic = "^2.10.1"
 click = "^8.1.7"
--- a/crates/common/src/configuration.rs
+++ b/crates/common/src/configuration.rs
@ -194,17 +194,11 @@ pub struct Parameter {
    pub in_path: Option<bool>,
 }

-#[derive(Debug, Clone, Serialize, Deserialize)]
-pub struct EndpointDetails {
-    pub name: String,
-    pub path: Option<String>,
-}
-
 #[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash, Default)]
 pub enum HttpMethod {
+    #[default]
    #[serde(rename = "GET")]
    Get,
-    #[default]
    #[serde(rename = "POST")]
    Post,
 }
@ -218,6 +212,14 @@ impl Display for HttpMethod {
    }
 }

+#[derive(Debug, Clone, Serialize, Deserialize)]
+pub struct EndpointDetails {
+    pub name: String,
+    pub path: Option<String>,
+    #[serde(rename = "http_method")]
+    pub method: Option<HttpMethod>,
+}
+
 #[derive(Debug, Clone, Serialize, Deserialize)]
 pub struct PromptTarget {
    pub name: String,
--- a/crates/prompt_gateway/src/stream_context.rs
+++ b/crates/prompt_gateway/src/stream_context.rs
@ -940,7 +940,7 @@ impl StreamContext {
        let endpoint = prompt_target.endpoint.unwrap();
        let path: String = endpoint.path.unwrap_or(String::from("/"));

-        // only add params that are of string and number type
+        // only add params that are of string, number and bool type
        let url_params = tool_params
            .iter()
            .filter(|(_, value)| value.is_number() || value.is_string() || value.is_bool())
@ -967,7 +967,7 @@ impl StreamContext {
            }
        };

-        let http_method = prompt_target.method.unwrap_or_default().to_string();
+        let http_method = endpoint.method.unwrap_or_default().to_string();
        let mut headers = vec![
            (ARCH_UPSTREAM_HOST_HEADER, endpoint.name.as_str()),
            (":method", &http_method),
--- a/crates/prompt_gateway/tests/integration.rs
+++ b/crates/prompt_gateway/tests/integration.rs
@ -360,6 +360,7 @@ prompt_targets:
    endpoint:
      name: api_server
      path: /weather
+      http_method: POST
    system_prompt: |
      You are a helpful weather forecaster. Use weater data that is provided to you. Please following following guidelines when responding to user queries:
      - Use farenheight for temperature
--- a/demos/currency_exchange/README.md
+++ b/demos/currency_exchange/README.md
@ -0,0 +1 @@
+This demo shows how you can use a publicly hosted rest api and interact it using arch gateway.
--- a/demos/currency_exchange/arch_config.yaml
+++ b/demos/currency_exchange/arch_config.yaml
@ -0,0 +1,52 @@
+version: v0.1
+
+listener:
+  address: 0.0.0.0
+  port: 10000
+  message_format: huggingface
+  connect_timeout: 0.005s
+
+llm_providers:
+  - name: gpt-4o
+    access_key: $OPENAI_API_KEY
+    provider: openai
+    model: gpt-4o
+
+system_prompt: |
+  You are a helpful assistant.
+
+prompt_guards:
+  input_guards:
+    jailbreak:
+      on_exception:
+        message: Looks like you're curious about my abilities, but I can only provide assistance for currency exchange.
+
+prompt_targets:
+  - name: currency_exchange
+    description: Get currency exchange rate from USD to other currencies
+    parameters:
+      - name: currency_symbol
+        description: the currency that needs conversion
+        required: true
+        type: str
+        in_path: true
+    endpoint:
+      name: frankfurther_api
+      path: /v1/latest?base=USD&symbols={currency_symbol}
+    system_prompt: |
+      You are a helpful assistant. Show me the currency symbol you want to convert from USD.
+
+  - name: get_supported_currencies
+    description: Get list of supported currencies for conversion
+    endpoint:
+      name: frankfurther_api
+      path: /v1/currencies
+
+endpoints:
+  frankfurther_api:
+    endpoint: api.frankfurter.dev:443
+    protocol: https
+
+tracing:
+  random_sampling: 100
+  trace_arch_internal: true
--- a/demos/currency_exchange/docker-compose.yaml
+++ b/demos/currency_exchange/docker-compose.yaml
@ -0,0 +1,21 @@
+services:
+  chatbot_ui:
+    build:
+      context: ../shared/chatbot_ui
+    ports:
+      - "18080:8080"
+    environment:
+      # this is only because we are running the sample app in the same docker container environemtn as archgw
+      - CHAT_COMPLETION_ENDPOINT=http://host.docker.internal:10000/v1
+    extra_hosts:
+      - "host.docker.internal:host-gateway"
+    volumes:
+      - ./arch_config.yaml:/app/arch_config.yaml
+
+  jaeger:
+    build:
+      context: ../shared/jaeger
+    ports:
+      - "16686:16686"
+      - "4317:4317"
+      - "4318:4318"
--- a/demos/weather_forecast_signoz/run_demo.sh
+++ b/demos/weather_forecast_signoz/run_demo.sh
--- a/demos/hr_agent/arch_config.yaml
+++ b/demos/hr_agent/arch_config.yaml
@ -34,6 +34,7 @@ prompt_targets:
      endpoint:
        name: app_server
        path: /agent/workforce
+        http_method: POST
      parameters:
        - name: staffing_type
          type: str
@ -52,6 +53,7 @@ prompt_targets:
      endpoint:
        name: app_server
        path: /agent/slack_message
+        http_method: POST
      description: sends a slack message on a channel
      parameters:
        - name: slack_message
--- a/demos/insurance_agent/arch_config.yaml
+++ b/demos/insurance_agent/arch_config.yaml
@ -29,6 +29,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /policy/qa
+      http_method: POST
    description: Handle general Q/A related to insurance.
    default: true

@ -37,6 +38,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /policy/coverage
+      http_method: POST
    parameters:
    - name: policy_type
      type: str
@ -48,6 +50,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /policy/initiate
+      http_method: POST
    description: Start a policy coverage for car, boat, motorcycle or house.
    parameters:
    - name: policy_type
@ -64,6 +67,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /policy/claim
+      http_method: POST
    description: Update the notes on the claim
    parameters:
    - name: claim_id
@ -79,6 +83,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /policy/deductible
+      http_method: POST
    description: Update the deductible amount for a specific policy coverage.
    parameters:
    - name: policy_id
--- a/demos/network_agent/arch_config.yaml
+++ b/demos/network_agent/arch_config.yaml
@ -22,6 +22,7 @@ prompt_targets:
      endpoint:
        name: app_server
        path: /agent/device_summary
+        http_method: POST
      parameters:
        - name: device_ids
          type: list
@ -36,6 +37,7 @@ prompt_targets:
      endpoint:
        name: app_server
        path: /agent/device_reboot
+        http_method: POST
      parameters:
        - name: device_ids
          type: list
--- a/demos/shared/logfire/Dockerfile
+++ b/demos/shared/logfire/Dockerfile
@ -0,0 +1,5 @@
+FROM otel/opentelemetry-collector:latest
+
+COPY otel-collector-config.yaml /etc/otel-collector-config.yaml
+
+ENTRYPOINT ["/otelcol", "--config=/etc/otel-collector-config.yaml"]
--- a/demos/shared/logfire/otel-collector-config.yaml
+++ b/demos/shared/logfire/otel-collector-config.yaml
@ -0,0 +1,24 @@
+receivers:
+  otlp:
+    protocols:
+      grpc:
+        endpoint: 0.0.0.0:4317
+      http:
+        endpoint: 0.0.0.0:4318
+
+exporters:
+  otlphttp:
+    endpoint: "https://logfire-api.pydantic.dev"
+    headers:
+      Authorization: "${LOGFIRE_API_KEY}"
+
+processors:
+  batch:
+    timeout: 5s
+
+service:
+  pipelines:
+    traces:
+      receivers: [otlp]
+      processors: [batch]
+      exporters: [otlphttp]
--- a/demos/shared/weather_forecast_service/Dockerfile
+++ b/demos/shared/weather_forecast_service/Dockerfile
--- a/demos/weather_forecast/README.md
+++ b/demos/weather_forecast/README.md
@ -1,27 +1,63 @@
 # Function calling
+
 This demo shows how you can use Arch's core function calling capabilites.

 # Starting the demo
+
 1. Please make sure the [pre-requisites](https://github.com/katanemo/arch/?tab=readme-ov-file#prerequisites) are installed correctly
 2. Start Arch

-3.
-   ```sh
+3. ```sh
   sh run_demo.sh
   ```
 4. Navigate to http://localhost:18080/
 5. You can type in queries like "how is the weather?"

 # Observability
+
 Arch gateway publishes stats endpoint at http://localhost:19901/stats. In this demo we are using prometheus to pull stats from arch and we are using grafana to visalize the stats in dashboard. To see grafana dashboard follow instructions below,

 1. Start grafana and prometheus using following command
   ```yaml
   docker compose --profile monitoring up
   ```
-1. Navigate to http://localhost:3000/ to open grafana UI (use admin/grafana as credentials)
-1. From grafana left nav click on dashboards and select "Intelligent Gateway Overview" to view arch gateway stats
-
+2. Navigate to http://localhost:3000/ to open grafana UI (use admin/grafana as credentials)
+3. From grafana left nav click on dashboards and select "Intelligent Gateway Overview" to view arch gateway stats

 Here is a sample interaction,
 <img width="575" alt="image" src="https://github.com/user-attachments/assets/e0929490-3eb2-4130-ae87-a732aea4d059">
+
+## Tracing
+
+To see a tracing dashboard follow instructions below,
+
+1. For Jaeger, you can either use the default run_demo.sh script or run the following command:
+
+```sh
+sh run_demo.sh jaeger
+```
+
+2. For Logfire, first make sure to add a LOGFIRE_API_KEY to the .env file. You can either use the default run_demo.sh script or run the following command:
+
+```sh
+sh run_demo.sh logfire
+```
+
+3. For Signoz, you can either use the default run_demo.sh script or run the following command:
+
+```sh
+sh run_demo.sh signoz
+```
+
+If using Jaeger, navigate to http://localhost:16686/ to open Jaeger UI
+
+If using Signoz, navigate to http://localhost:3301/ to open Signoz UI
+
+If using Logfire, navigate to your logfire dashboard that you got the write key from to view the dashboard
+
+### Stopping Demo
+
+1. To end the demo, run the following command:
+   ```sh
+   sh run_demo.sh down
+   ```
--- a/demos/weather_forecast/arch_config.yaml
+++ b/demos/weather_forecast/arch_config.yaml
@ -60,6 +60,7 @@ prompt_targets:
    endpoint:
      name: weather_forecast_service
      path: /weather
+      http_method: POST

  - name: default_target
    default: true
@ -67,6 +68,7 @@ prompt_targets:
    endpoint:
      name: weather_forecast_service
      path: /default_target
+      http_method: POST
    system_prompt: |
      You are a helpful assistant! Summarize the user's request and provide a helpful response.
    # if it is set to false arch will send response that it received from this prompt target to the user
--- a/demos/weather_forecast/docker-compose-jaeger.yaml
+++ b/demos/weather_forecast/docker-compose-jaeger.yaml
@ -0,0 +1,41 @@
+services:
+  weather_forecast_service:
+    build:
+      context: ./
+    environment:
+      - OLTP_HOST=http://jaeger:4317
+    extra_hosts:
+      - "host.docker.internal:host-gateway"
+    ports:
+      - "18083:80"
+
+  chatbot_ui:
+    build:
+      context: ../shared/chatbot_ui
+    ports:
+      - "18080:8080"
+    environment:
+      # this is only because we are running the sample app in the same docker container environemtn as archgw
+      - CHAT_COMPLETION_ENDPOINT=http://host.docker.internal:10000/v1
+    extra_hosts:
+      - "host.docker.internal:host-gateway"
+    volumes:
+      - ./arch_config.yaml:/app/arch_config.yaml
+
+  jaeger:
+    build:
+      context: ../shared/jaeger
+    ports:
+      - "16686:16686"
+      - "4317:4317"
+      - "4318:4318"
+
+  prometheus:
+    build:
+      context: ../shared/prometheus
+
+  grafana:
+    build:
+      context: ../shared/grafana
+    ports:
+      - "3000:3000"
--- a/demos/weather_forecast/docker-compose-logfire.yaml
+++ b/demos/weather_forecast/docker-compose-logfire.yaml
@ -0,0 +1,46 @@
+services:
+  weather_forecast_service:
+    build:
+      context: ./weather_forecast_service
+    environment:
+      - OLTP_HOST=http://otel-collector:4317
+    extra_hosts:
+      - "host.docker.internal:host-gateway"
+    ports:
+      - "18083:80"
+
+  chatbot_ui:
+    build:
+      context: ../shared/chatbot_ui
+    ports:
+      - "18080:8080"
+    environment:
+      # this is only because we are running the sample app in the same docker container environment as archgw
+      - CHAT_COMPLETION_ENDPOINT=http://host.docker.internal:10000/v1
+    extra_hosts:
+      - "host.docker.internal:host-gateway"
+    volumes:
+      - ./arch_config.yaml:/app/arch_config.yaml
+
+  otel-collector:
+    build:
+      context: ../shared/logfire/
+    ports:
+      - "4317:4317"
+      - "4318:4318"
+    volumes:
+      - ../shared/logfire/otel-collector-config.yaml:/etc/otel-collector-config.yaml
+    env_file:
+      - .env
+    environment:
+      - LOGFIRE_API_KEY
+
+  prometheus:
+    build:
+      context: ../shared/prometheus
+
+  grafana:
+    build:
+      context: ../shared/grafana
+    ports:
+      - "3000:3000"
--- a/demos/weather_forecast/docker-compose-signoz.yaml
+++ b/demos/weather_forecast/docker-compose-signoz.yaml
@ -4,7 +4,7 @@ include:
 services:
  weather_forecast_service:
    build:
-      context: ../shared/weather_forecast_service
+      context: ./weather_forecast_service
    environment:
      - OLTP_HOST=http://otel-collector:4317
    extra_hosts:
--- a/demos/weather_forecast/docker-compose.yaml
+++ b/demos/weather_forecast/docker-compose.yaml
@ -1,7 +1,7 @@
 services:
  weather_forecast_service:
    build:
-      context: ../shared/weather_forecast_service
+      context: ./
    environment:
      - OLTP_HOST=http://jaeger:4317
    extra_hosts:
--- a/demos/shared/weather_forecast_service/main.py
+++ b/demos/shared/weather_forecast_service/main.py
--- a/demos/shared/weather_forecast_service/poetry.lock
+++ b/demos/shared/weather_forecast_service/poetry.lock
--- a/demos/shared/weather_forecast_service/pyproject.toml
+++ b/demos/shared/weather_forecast_service/pyproject.toml
--- a/demos/weather_forecast/run_demo.sh
+++ b/demos/weather_forecast/run_demo.sh
@ -1,47 +1,95 @@
 #!/bin/bash
 set -e

+# Function to load environment variables from the .env file
+load_env() {
+  if [ -f ".env" ]; then
+    export $(grep -v '^#' .env | xargs)
+  fi
+}
+
+# Function to determine the docker-compose file based on the argument
+get_compose_file() {
+  case "$1" in
+    jaeger)
+      echo "docker-compose-jaeger.yaml"
+      ;;
+    logfire)
+      echo "docker-compose-logfire.yaml"
+      ;;
+    signoz)
+      echo "docker-compose-signoz.yaml"
+      ;;
+    *)
+      echo "docker-compose.yaml"
+      ;;
+  esac
+}
+
 # Function to start the demo
 start_demo() {
-  # Step 1: Check if .env file exists
+  # Step 1: Determine the docker-compose file
+  COMPOSE_FILE=$(get_compose_file "$1" 2>/dev/null)
+
+  # Step 2: Check if .env file exists
  if [ -f ".env" ]; then
    echo ".env file already exists. Skipping creation."
  else
-    # Step 2: Create `.env` file and set OpenAI key
+    # Step 3: Check for required environment variables
    if [ -z "$OPENAI_API_KEY" ]; then
      echo "Error: OPENAI_API_KEY environment variable is not set for the demo."
      exit 1
    fi
+    if [ "$1" == "logfire" ] && [ -z "$LOGFIRE_API_KEY" ]; then
+      echo "Error: LOGFIRE_API_KEY environment variable is required for Logfire."
+      exit 1
+    fi

+    # Create .env file
    echo "Creating .env file..."
    echo "OPENAI_API_KEY=$OPENAI_API_KEY" > .env
-    echo ".env file created with OPENAI_API_KEY."
+    if [ "$1" == "logfire" ]; then
+      echo "LOGFIRE_API_KEY=$LOGFIRE_API_KEY" >> .env
+    fi
+    echo ".env file created with required API keys."
  fi

-  # Step 3: Start Arch
+  load_env
+
+  if [ "$1" == "logfire" ] && [ -z "$LOGFIRE_API_KEY"]; then
+    echo "Error: LOGFIRE_API_KEY environment variable is required for Logfire."
+    exit 1
+  fi
+
+  # Step 4: Start Arch
  echo "Starting Arch with arch_config.yaml..."
  archgw up arch_config.yaml

-  # Step 4: Start Network Agent
-  echo "Starting Network Agent using Docker Compose..."
-  docker compose up -d  # Run in detached mode
+  # Step 5: Start Network Agent with the chosen Docker Compose file
+  echo "Starting Network Agent with $COMPOSE_FILE..."
+  docker compose -f "$COMPOSE_FILE" up -d  # Run in detached mode
 }

 # Function to stop the demo
 stop_demo() {
-  # Step 1: Stop Docker Compose services
-  echo "Stopping Network Agent using Docker Compose..."
-  docker compose down
+  echo "Stopping all Docker Compose services..."

-  # Step 2: Stop Arch
+  # Stop all services by iterating through all configurations
+  for compose_file in ./docker-compose*.yaml; do
+    echo "Stopping services in $compose_file..."
+    docker compose -f "$compose_file" down
+  done
+
+  # Stop Arch
  echo "Stopping Arch..."
  archgw down
 }

 # Main script logic
 if [ "$1" == "down" ]; then
+  # Call stop_demo with the second argument as the demo to stop
  stop_demo
 else
-  # Default action is to bring the demo up
-  start_demo
+  # Use the argument (jaeger, logfire, signoz) to determine the compose file
+  start_demo "$1"
 fi
--- a/demos/weather_forecast_signoz/README.md
+++ b/demos/weather_forecast_signoz/README.md
@ -1,29 +0,0 @@
-# Function calling
-This demo shows how you can use Arch's core function calling capabilites.
-
-# Starting the demo
-1. Please make sure the [pre-requisites](https://github.com/katanemo/arch/?tab=readme-ov-file#prerequisites) are installed correctly
-2. Start Arch
-
-3.
-   ```sh
-   sh run_demo.sh
-   ```
-4. Navigate to http://localhost:18080/
-5. You can type in queries like "how is the weather?"
-
-# Observability
-Arch gateway publishes stats endpoint at http://localhost:19901/stats. In this demo we are using prometheus to pull stats from arch and we are using grafana to visalize the stats in dashboard. To see grafana dashboard follow instructions below,
-
-1. Start grafana and prometheus using following command
-   ```yaml
-   docker compose --profile monitoring up
-   ```
-1. Navigate to http://localhost:3000/ to open grafana UI (use admin/grafana as credentials)
-1. From grafana left nav click on dashboards and select "Intelligent Gateway Overview" to view arch gateway stats
-
-
-Here is a sample interaction,
-<img width="575" alt="image" src="https://github.com/user-attachments/assets/e0929490-3eb2-4130-ae87-a732aea4d059">
-
-1. Signoz UI: http://localhost:3301
--- a/demos/weather_forecast_signoz/arch_config.yaml
+++ b/demos/weather_forecast_signoz/arch_config.yaml
@ -1,72 +0,0 @@
-version: "0.1-beta"
-
-listener:
-  address: 0.0.0.0
-  port: 10000
-  message_format: huggingface
-  connect_timeout: 0.005s
-
-endpoints:
-  weather_forecast_service:
-    endpoint: host.docker.internal:18083
-    connect_timeout: 0.005s
-
-overrides:
-  # confidence threshold for prompt target intent matching
-  prompt_target_intent_matching_threshold: 0.6
-
-llm_providers:
-  - name: gpt-4o-mini
-    access_key: $OPENAI_API_KEY
-    provider: openai
-    model: gpt-4o-mini
-    default: true
-
-  - name: gpt-3.5-turbo-0125
-    access_key: $OPENAI_API_KEY
-    provider: openai
-    model: gpt-3.5-turbo-0125
-
-  - name: gpt-4o
-    access_key: $OPENAI_API_KEY
-    provider: openai
-    model: gpt-4o
-
-system_prompt: |
-  You are a helpful assistant.
-
-prompt_targets:
-  - name: weather_forecast
-    description: Check weather information for a given city.
-    parameters:
-      - name: city
-        description: the name of the city
-        required: true
-        type: str
-      - name: days
-        description: the number of days
-        type: int
-        required: true
-      - name: units
-        description: the temperature unit, e.g., Celsius and Fahrenheit
-        type: str
-        default: Fahrenheit
-    endpoint:
-      name: weather_forecast_service
-      path: /weather
-
-  - name: default_target
-    default: true
-    description: This is the default target for all unmatched prompts.
-    endpoint:
-      name: weather_forecast_service
-      path: /default_target
-    system_prompt: |
-      You are a helpful assistant! Summarize the user's request and provide a helpful response.
-    # if it is set to false arch will send response that it received from this prompt target to the user
-    # if true arch will forward the response to the default LLM
-    auto_llm_dispatch_on_response: false
-
-tracing:
-  random_sampling: 100
-  # trace_arch: true
--- a/docs/source/get_started/includes/quickstart.yaml
+++ b/docs/source/get_started/includes/quickstart.yaml
@ -1,43 +0,0 @@
-version: v0.1
-listener:
-  address: 127.0.0.1
-  port: 8080 #If you configure port 443, you'll need to update the listener with tls_certificates
-  message_format: huggingface
-
-# Centralized way to manage LLMs, manage keys, retry logic, failover and limits in a central way
-llm_providers:
-  - name: OpenAI
-    provider: openai
-    access_key: $OPENAI_API_KEY
-    model: gpt-3.5-turbo
-    default: true
-
-# default system prompt used by all prompt targets
-system_prompt: |
-  You are a network assistant that helps operators with a better understanding of network traffic flow and perform actions on networking operations. No advice on manufacturers or purchasing decisions.
-
-prompt_targets:
-    - name: device_summary
-      description: Retrieve network statistics for specific devices within a time range
-      endpoint:
-        name: app_server
-        path: /agent/device_summary
-      parameters:
-        - name: device_ids
-          type: list
-          description: A list of device identifiers (IDs) to retrieve statistics for.
-          required: true  # device_ids are required to get device statistics
-        - name: days
-          type: int
-          description: The number of days for which to gather device statistics.
-          default: "7"
-
-# Arch creates a round-robin load balancing between different endpoints, managed via the cluster subsystem.
-endpoints:
-  app_server:
-    # value could be ip address or a hostname with port
-    # this could also be a list of endpoints for load balancing
-    # for example endpoint: [ ip1:port, ip2:port ]
-    endpoint: host.docker.internal:18083
-    # max time to wait for a connection to be established
-    connect_timeout: 0.005s
--- a/docs/source/get_started/quickstart.rst
+++ b/docs/source/get_started/quickstart.rst
@ -7,74 +7,255 @@ Follow this guide to learn how to quickly set up Arch and integrate it into your


 Prerequisites
----------------------------
+-------------

 Before you begin, ensure you have the following:

-.. vale Vale.Spelling = NO
+1. `Docker System <https://docs.docker.com/get-started/get-docker/>`_ (v24)
+2. `Docker compose <https://docs.docker.com/compose/install/>`_ (v2.29)
+3. `Python <https://www.python.org/downloads/>`_ (v3.12)

- ``Docker`` & ``Python`` installed on your system
- ``API Keys`` for LLM providers (if using external LLMs)
-
-The fastest way to get started using Arch is to use `katanemo/archgw <https://hub.docker.com/r/katanemo/archgw>`_ pre-built binaries.
-You can also build it from source.
-
-
-Step 1: Install Arch
----------------------------
-Arch's CLI allows you to manage and interact with the Arch gateway efficiently. To install the CLI, simply
-run the following command:
-
-.. code-block:: console
-
-    $ pip install archgw
-
-This will install the archgw command-line tool globally on your system.
+Arch's CLI allows you to manage and interact with the Arch gateway efficiently. To install the CLI, simply run the following command:

 .. tip::
-   We recommend that developers create a new Python virtual environment to isolate dependencies before installing Arch.
-   This ensures that `archgw` and its dependencies do not interfere with other packages on your system.

-   To create and activate a virtual environment, you can run the following commands:
-
-   .. code-block:: console
-
-      $ python -m venv venv
-      $ source venv/bin/activate   # On Windows, use: venv\Scripts\activate
-      $ pip install archgw
-
-
-Step 2: Config Arch
-------------------
-
-Arch operates based on a configuration file where you can define LLM providers, prompt targets, and guardrails, etc.
-Below is an example configuration to get you started, including:
-
-.. vale Vale.Spelling = NO
-
- ``endpoints``: Specifies where Arch listens for incoming prompts.
- ``system_prompts``: Defines predefined prompts to set the context for interactions.
- ``llm_providers``: Lists the LLM providers Arch can route prompts to.
- ``prompt_guards``: Sets up rules to detect and reject undesirable prompts.
- ``prompt_targets``: Defines endpoints that handle specific types of prompts.
- ``error_target``: Specifies where to route errors for handling.
-
-.. literalinclude:: includes/quickstart.yaml
-    :language: yaml
-
-
-Step 3: Start Arch Gateway
--------------------------
+   We recommend that developers create a new Python virtual environment to isolate dependencies before installing Arch. This ensures that ``archgw`` and its dependencies do not interfere with other packages on your system.

 .. code-block:: console

-    $ archgw up [path_to_config]
+   $ python -m venv venv
+   $ source venv/bin/activate   # On Windows, use: venv\Scripts\activate
+   $ pip install archgw==0.1.6

-For detailed usage please refer to the `Support <https://github.com/katanemo/arch/blob/main/arch/tools/README.md#setup-instructionsuser-archgw-cli>`
+
+Build AI Agent with Arch Gateway
+--------------------------------
+
+In the following quickstart, we will show you how easy it is to build an AI agent with the Arch gateway. We will build a currency exchange agent using the following simple steps. For this demo, we will use `https://api.frankfurter.dev/` to fetch the latest prices for currencies and assume USD as the base currency.
+
+Step 1. Create arch config file
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+Create ``arch_config.yaml`` file with the following content:
+
+.. code-block:: yaml
+
+   version: v0.1
+
+   listener:
+     address: 0.0.0.0
+     port: 10000
+     message_format: huggingface
+     connect_timeout: 0.005s
+
+   llm_providers:
+     - name: gpt-4o
+       access_key: $OPENAI_API_KEY
+       provider: openai
+       model: gpt-4o
+
+   system_prompt: |
+     You are a helpful assistant.
+
+   prompt_guards:
+     input_guards:
+       jailbreak:
+         on_exception:
+           message: Looks like you're curious about my abilities, but I can only provide assistance for currency exchange.
+
+   prompt_targets:
+     - name: currency_exchange
+       description: Get currency exchange rate from USD to other currencies
+       parameters:
+         - name: currency_symbol
+           description: the currency that needs conversion
+           required: true
+           type: str
+           in_path: true
+       endpoint:
+         name: frankfurther_api
+         path: /v1/latest?base=USD&symbols={currency_symbol}
+       system_prompt: |
+         You are a helpful assistant. Show me the currency symbol you want to convert from USD.
+
+     - name: get_supported_currencies
+       description: Get list of supported currencies for conversion
+       endpoint:
+         name: frankfurther_api
+         path: /v1/currencies
+
+   endpoints:
+     frankfurther_api:
+       endpoint: api.frankfurter.dev:443
+       protocol: https
+
+Step 2. Start arch gateway with currency conversion config
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+.. code-block:: sh
+
+   $ archgw up arch_config.yaml
+   2024-12-05 16:56:27,979 - cli.main - INFO - Starting archgw cli version: 0.1.5
+   ...
+   2024-12-05 16:56:28,485 - cli.utils - INFO - Schema validation successful!
+   2024-12-05 16:56:28,485 - cli.main - INFO - Starting arch model server and arch gateway
+   ...
+   2024-12-05 16:56:51,647 - cli.core - INFO - Container is healthy!
+
+Once the gateway is up, you can start interacting with it at port 10000 using the OpenAI chat completion API.
+
+Some sample queries you can ask include: ``what is currency rate for gbp?`` or ``show me list of currencies for conversion``.
+
+Step 3. Interacting with gateway using curl command
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+Here is a sample curl command you can use to interact:
+
+.. code-block:: bash
+
+   $ curl --header 'Content-Type: application/json' \
+     --data '{"messages": [{"role": "user","content": "what is exchange rate for gbp"}]}' \
+     http://localhost:10000/v1/chat/completions | jq ".choices[0].message.content"
+
+   "As of the date provided in your context, December 5, 2024, the exchange rate for GBP (British Pound) from USD (United States Dollar) is 0.78558. This means that 1 USD is equivalent to 0.78558 GBP."
+
+And to get the list of supported currencies:
+
+.. code-block:: bash
+
+   $ curl --header 'Content-Type: application/json' \
+     --data '{"messages": [{"role": "user","content": "show me list of currencies that are supported for conversion"}]}' \
+     http://localhost:10000/v1/chat/completions | jq ".choices[0].message.content"
+
+   "Here is a list of the currencies that are supported for conversion from USD, along with their symbols:\n\n1. AUD - Australian Dollar\n2. BGN - Bulgarian Lev\n3. BRL - Brazilian Real\n4. CAD - Canadian Dollar\n5. CHF - Swiss Franc\n6. CNY - Chinese Renminbi Yuan\n7. CZK - Czech Koruna\n8. DKK - Danish Krone\n9. EUR - Euro\n10. GBP - British Pound\n11. HKD - Hong Kong Dollar\n12. HUF - Hungarian Forint\n13. IDR - Indonesian Rupiah\n14. ILS - Israeli New Sheqel\n15. INR - Indian Rupee\n16. ISK - Icelandic Króna\n17. JPY - Japanese Yen\n18. KRW - South Korean Won\n19. MXN - Mexican Peso\n20. MYR - Malaysian Ringgit\n21. NOK - Norwegian Krone\n22. NZD - New Zealand Dollar\n23. PHP - Philippine Peso\n24. PLN - Polish Złoty\n25. RON - Romanian Leu\n26. SEK - Swedish Krona\n27. SGD - Singapore Dollar\n28. THB - Thai Baht\n29. TRY - Turkish Lira\n30. USD - United States Dollar\n31. ZAR - South African Rand\n\nIf you want to convert USD to any of these currencies, you can select the one you are interested in."
+
+
+Use Arch Gateway as LLM Router
+------------------------------
+
+Step 1. Create arch config file
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+Arch operates based on a configuration file where you can define LLM providers, prompt targets, guardrails, etc. Below is an example configuration that defines OpenAI and Mistral LLM providers.
+
+Create ``arch_config.yaml`` file with the following content:
+
+.. code-block:: yaml
+
+   version: v0.1
+
+   listener:
+     address: 0.0.0.0
+     port: 10000
+     message_format: huggingface
+     connect_timeout: 0.005s
+
+   llm_providers:
+     - name: gpt-4o
+       access_key: $OPENAI_API_KEY
+       provider: openai
+       model: gpt-4o
+       default: true
+
+     - name: ministral-3b
+       access_key: $MISTRAL_API_KEY
+       provider: mistral
+       model: ministral-3b-latest
+
+Step 2. Start arch gateway
+~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+Once the config file is created, ensure that you have environment variables set up for ``MISTRAL_API_KEY`` and ``OPENAI_API_KEY`` (or these are defined in a ``.env`` file).
+
+Start the Arch gateway:
+
+.. code-block:: console
+
+   $ archgw up arch_config.yaml
+   2024-12-05 11:24:51,288 - cli.main - INFO - Starting archgw cli version: 0.1.5
+   2024-12-05 11:24:51,825 - cli.utils - INFO - Schema validation successful!
+   2024-12-05 11:24:51,825 - cli.main - INFO - Starting arch model server and arch gateway
+   ...
+   2024-12-05 11:25:16,131 - cli.core - INFO - Container is healthy!
+
+Step 3: Interact with LLM
+~~~~~~~~~~~~~~~~~~~~~~~~~
+
+Step 3.1: Using OpenAI Python client
++++++++++++++++++++++++++++++++++++
+
+Make outbound calls via the Arch gateway:
+
+.. code-block:: python
+
+   from openai import OpenAI
+
+   # Use the OpenAI client as usual
+   client = OpenAI(
+     # No need to set a specific openai.api_key since it's configured in Arch's gateway
+     api_key='--',
+     # Set the OpenAI API base URL to the Arch gateway endpoint
+     base_url="http://127.0.0.1:12000/v1"
+   )
+
+   response = client.chat.completions.create(
+       # we select model from arch_config file
+       model="--",
+       messages=[{"role": "user", "content": "What is the capital of France?"}],
+   )
+
+   print("OpenAI Response:", response.choices[0].message.content)
+
+Step 3.2: Using curl command
++++++++++++++++++++++++++++
+
+.. code-block:: bash
+
+   $ curl --header 'Content-Type: application/json' \
+     --data '{"messages": [{"role": "user","content": "What is the capital of France?"}]}' \
+     http://localhost:12000/v1/chat/completions
+
+   {
+     ...
+     "model": "gpt-4o-2024-08-06",
+     "choices": [
+       {
+         ...
+         "message": {
+           "role": "assistant",
+           "content": "The capital of France is Paris.",
+         },
+       }
+     ],
+   }
+
+You can override model selection using the ``x-arch-llm-provider-hint`` header. For example, to use Mistral, use the following curl command:
+
+.. code-block:: bash
+
+   $ curl --header 'Content-Type: application/json' \
+     --header 'x-arch-llm-provider-hint: ministral-3b' \
+     --data '{"messages": [{"role": "user","content": "What is the capital of France?"}]}' \
+     http://localhost:12000/v1/chat/completions
+
+   {
+     ...
+     "model": "ministral-3b-latest",
+     "choices": [
+       {
+         "message": {
+           "role": "assistant",
+           "content": "The capital of France is Paris. It is the most populous city in France and is known for its iconic landmarks such as the Eiffel Tower, the Louvre Museum, and Notre-Dame Cathedral. Paris is also a major global center for art, fashion, gastronomy, and culture.",
+         },
+         ...
+       }
+     ],
+     ...
+   }


 Next Steps
-------------------
+==========

 Congratulations! You've successfully set up Arch and made your first prompt-based request. To further enhance your GenAI applications, explore the following resources:

--- a/docs/source/guides/includes/arch_config.yaml
+++ b/docs/source/guides/includes/arch_config.yaml
@ -31,6 +31,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /agent/summary
+      http_method: POST
    # Arch uses the default LLM and treats the response from the endpoint as the prompt to send to the LLM
    auto_llm_dispatch_on_response: true
    # override system prompt for this prompt target
@ -41,6 +42,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /agent/action
+      http_method: POST
    parameters:
      - name: device_id
        type: str
--- a/docs/source/resources/includes/arch_config_full_reference.yaml
+++ b/docs/source/resources/includes/arch_config_full_reference.yaml
@ -77,6 +77,7 @@ prompt_targets:
    endpoint:
      name: app_server
      path: /agent/summary
+      http_method: POST
    # Arch uses the default LLM and treats the response from the endpoint as the prompt to send to the LLM
    auto_llm_dispatch_on_response: true
    # override system prompt for this prompt target
--- a/e2e_tests/run_e2e_tests.sh
+++ b/e2e_tests/run_e2e_tests.sh
@ -39,7 +39,7 @@ cd -
 log building and installing archgw cli
 log ==================================
 cd ../arch/tools
-sh build_cli.sh
+poetry install
 cd -

 log building docker image for arch gateway
--- a/model_server/poetry.lock
+++ b/model_server/poetry.lock
--- a/model_server/pyproject.toml
+++ b/model_server/pyproject.toml
@ -1,6 +1,6 @@
 [tool.poetry]
 name = "archgw_modelserver"
-version = "0.1.5"
+version = "0.1.6"
 description = "A model server for serving models"
 authors = ["Katanemo Labs, Inc <archgw@katanemo.com>"]
 license = "Apache 2.0"
@ -29,7 +29,7 @@ openai = "1.50.2"
 tf-keras = "*"
 onnx = "1.17.0"
 onnxruntime = "1.19.2"
-httpx = "*"
+httpx = "0.27.2" # https://community.openai.com/t/typeerror-asyncclient-init-got-an-unexpected-keyword-argument-proxies/1040287
 pytest-asyncio = "*"
 pytest = "*"
 opentelemetry-api = "^1.28.0"
				`@ -0,0 +1 @@`
				`This demo shows how you can use a publicly hosted rest api and interact it using arch gateway.`