The duration and tokenUsage numberRange filters cast the JSON text value
to INTEGER, but some rows store call_duration_seconds as a JSON float
(0.0, written by get_call_duration when the pipeline never recorded a
start time). Postgres cannot cast the text '0.0' to integer, so any
duration filter on /organizations/usage/runs failed with
InvalidTextRepresentationError. Cast to FLOAT instead, matching the
duration sort clause, and make get_call_duration return an int so new
rows are written as integers.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs: flesh out stub node pages for Global, Start Call, Agent, End Call, QA
Each page was previously a stub (1 sentence or a single <Note>). Now contains:
- Plain-English explanation of what the node does and when it runs
- Fields/configuration reference
- Common mistakes section
- Next Steps with CardGroup links
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: fix heading structure, Warning placement, variable formatting in node pages
- qa.mdx: promote Viewing QA Results from h3 to h2 (was incorrectly nested under Where it sits on the canvas)
- end-call.mdx: move Warning inside Common Mistakes, directly after the bullet it relates to
- start-call/agent/end-call/qa.mdx: wrap bare gathered_context and initial_context in backticks in Card descriptions
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: correct node page errors found by cross-checking against live API + DTO
start-call:
- Fix Node Name description (canvas label, not agent spoken name)
- Split Greeting Text and Prompt into separate documented fields (were conflated)
- Add Delayed Start field (real config option, was missing entirely)
- Fix Common Mistakes (remove wrong Agent Name claim)
- Add Pre-recorded Audio note with link
end-call:
- Remove incorrect <Note> claiming one End Call per workflow
- DTO has no max_instances constraint; llm_hint explicitly says multiple are supported
- Fix Common Mistakes to reflect correct guidance (multiple allowed)
agent:
- Add Extracting variables section documenting extraction_enabled, extraction_prompt,
extraction_variables fields with a concrete example
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: fix field names and QA model against live platform schemas
- qa.mdx: replace fictional criteria-list concept with accurate
qa_system_prompt field; document default output format (tags, score,
sentiment, summary); add sampling/duration/voicemail filter settings
- global.mdx: correct prompt-merge order (global prepended, not appended)
- start-call.mdx: update stale "Greeting Prompt" reference in common
mistakes to match renamed field "Greeting Text"
Verified against live node schemas via Dograh MCP get_node_type.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: fix em dash violations and appended/prepended contradiction
- global.mdx: "silently appended" contradicted "global content first";
changed to "prepended" for internal consistency and schema accuracy
- qa.mdx: remove two em dashes (hard no per style rules)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: correct node docs against live workflow SDK evidence
- start-call: replace invented fixed edge labels with accurate explanation
that edge labels and conditions are user-defined; add real examples
- end-call: rename "Final Message" to "Prompt" to match SDK field name
used in all real workflows; update all references
- global: add Start Call to Add Global Prompt toggle coverage
(startCall has add_global_prompt: true by default per schema)
- qa: replace informal tag descriptions with actual default tag names
from the qa_system_prompt schema
All changes verified against live Dograh MCP node schemas and
real workflow SDK code (workflow IDs 7764, 8291).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: flesh out all 5 voice agent node pages with accurate UI fields
Complete rewrite of stub docs for global, start-call, agent, end-call,
and qa nodes. Field names, toggle labels, and section headings now match
the actual Dograh UI exactly. Covers all fields visible in each node
edit panel including Allow Interruption, Add Global Prompt, Delayed Start,
Enable Variable Extraction, Pre-Call Data Fetch, and QA settings table.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: remove Allow Interruption from End Call node
End Call does not expose an Allow Interruption toggle in the UI.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* docs: address cubic and greptile review comments on node pages
- agent.mdx, end-call.mdx: remove/correct gathered_context claims -
it is not available in any node's Prompt, only in Webhook payloads
and the run record (confirmed against context-and-variables.mdx)
- global.mdx: fix appended -> prepended in frontmatter description
and Common Mistakes (contradicted How It Works section)
- introduction.mdx: update End Call node-type table - multiple End
Call nodes are supported, not just one
- agent.mdx, end-call.mdx, qa.mdx, start-call.mdx: tag bare code
fences as text for Mintlify syntax highlighting
- global.mdx, start-call.mdx, agent.mdx, end-call.mdx, qa.mdx:
convert absolute internal links to relative paths per style guide
* docs: embed video tutorials on all 5 node pages, fix cubic grammar comment
- Add Video Tutorial section (iframe embed) to global, start-call,
agent, end-call, and qa node pages, matching the pattern already
used in webhook.mdx
- global.mdx: fix "Global Node contain" -> "Global Node contains" in
frontmatter description, per cubic review on 09c5ff31
* docs: fix appended/prepended in introduction.mdx node table
Global node's node-type summary still said the prompt is "appended"
after Add Global Prompt is enabled, contradicting the corrected
"prepended" behavior documented everywhere else. Caught by greptile
P1 review on commit b7fe2d0d.
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: save call metadata in gathered context for an api trigger outbound call
* fix: handle workflow run update failures
* fix: format
* fix: pin integration workflow runs to published definitions
* test: expect text chat runtime configuration
* fix: move draft and template context handling out of create_workflow_run
create_workflow_run now only creates
the run with the definition_id and initial_context provided by the caller. Test call paths explicitly resolve draft definitions
and merge template context variables before creating the run, while production/runtime paths bind the published definition
without adding template defaults.
* fix: review comments
* chore: format and minor cleanups
---------
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>
* feat(tts): add websocket transport option for xAI TTS
xAI's pipecat service ships two implementations: XAIHttpTTSService
(batch REST, used by the existing xAI integration) and XAITTSService
(realtime WebSocket streaming). This adds a transport field on
XAITTSConfiguration ("http", default, unchanged behavior | "websocket")
and wires the latter into create_tts_service(), so live voice calls can
opt into lower time-to-first-byte streaming instead of request/response.
Originally built and battle-tested independently against v1.41.0 in a
production deployment (7 live Telnyx calls, 2026-07-15) before this PR
existed; ported onto current main and given a config toggle instead of
a second provider entry so it composes with the existing xAI HTTP path
rather than duplicating it.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* refactor(tts): use websocket-only xAI TTS, drop the transport toggle
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Sabiha Khan <sabihak89@gmail.com>
* ViciDial Working
* feat(telephony): add FreeSWITCH provider to upstream-PBX seam
Generalize the upstream-PBX capture and control paths to dispatch by
provider. FreeSWITCH bridges in with X-PBX-* headers and is driven over
the Event Socket Library (uuid_kill / uuid_transfer) by the channel UUID,
alongside the existing VICIdial ra_call_control path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* add vici specific configs
* feat(telephony): env-based VICIdial config + address3 post-call routing
Move VICIdial agent/non-agent API and FreeSWITCH ESL connection settings
out of hardcoded POC values into environment variables so the same image
works against a PBX on another server.
Add VICIdial update_lead forwarding: X-VICI-UPDATE-LEAD_* extracted
variables are mapped to lead columns and pushed via the non-agent API
before a transfer, with reserved API-control params dropped.
Add hardcoded address3 disposition routing (Y/N -> in-group transfer,
else hang up) and a synchronous final variable extraction so the
transfer path sees the freshest conversation state. Skip upstream_pbx
capture on non-PJSIP channels to avoid Asterisk 500s.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* chore: run formatter
* feat: make vici configurable from UI
* chore: clean up PR
* fix: incorporate review comments
* chore: incorporate review comments
* chore: generate client
* chore: incporporate review comments
* chore: incorporate review comments
---------
Co-authored-by: Dograh POC <payment@dograh.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(api): validate pagination bounds on run-list endpoints (#553)
GET /workflow/{id}/runs and GET /campaign/{id}/runs declared bare
`page: int = 1` / `limit: int = 50` params, then computed
`total_pages = (total_count + limit - 1) // limit`. A `?limit=0` raised an
unhandled ZeroDivisionError (HTTP 500), and negative limit/page produced a
negative offset and nonsensical pagination.
Add `Query(ge=1, le=100)` / `Query(1, ge=1)` bounds to both endpoints,
matching the sibling list endpoints (/usage/runs and the superuser runs
endpoint) that already validate these. Out-of-range values now return 422.
Adds a regression test covering limit=0/-5/101 and page=0 on both endpoints.
Fixes#553
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* chore(docs): regenerate openapi.json for run-list pagination bounds (#553)
The added Query(ge/le) bounds on the workflow-run and campaign-run list
endpoints changed the OpenAPI schema; regenerate the committed spec via
`python -m scripts.dump_docs_openapi` so the drift-check passes. Only the
limit/page parameter schemas for those two endpoints change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>
The OpenAI realtime factory already reads `language` from the realtime
config, but builds `InputAudioTranscription()` without it, so the language
is silently dropped and the model auto-detects on every utterance.
On 8kHz telephony audio this misfires badly: in our production tests a
Brazilian Portuguese speaker was transcribed as English, French and Chinese
within a single call, which then corrupted downstream extraction.
Every other STT branch in `create_stt_service` already honours
`language` (Deepgram, Google, Cartesia, Dograh, Sarvam), and
`GoogleRealtimeLLMConfiguration` already exposes a `language` field for
Gemini Live. This brings the OpenAI realtime provider in line with both.
- expose `language` on `OpenAIRealtimeLLMConfiguration` (optional,
defaults to None -> unchanged auto-detect behaviour)
- pass it through to `InputAudioTranscription`, which already accepts it
Verified against pipecat: the session now carries
`{"transcription": {"model": "gpt-realtime-whisper", "language": "pt"}}`.
Co-authored-by: Liberty Card <tecnologialibertycard@gmail.com>
* feat: add tool test panel for HTTP API tools
Lets developers run a saved HTTP API tool against its real endpoint
from the tool detail page, without needing a live call. Reuses the
production execute_http_tool path so test behavior matches call-time
behavior.
- New POST /tools/{tool_uuid}/test route
- Test panel with per-parameter typed inputs and auto-detected
context variable inputs (from preset parameter templates)
- Validate parameter name uniqueness on save, matching the existing
transferParameters check
- Fix stale FunctionCallsFromLLMInfoFrame import causing test
collection failures against pipecat-ai 1.5.0
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: revert local-venv pipecat drift fix, apply ruff import formatting
test_custom_tools.py and test_unregistered_function_call.py were edited
locally to drop FunctionCallsFromLLMInfoFrame after hitting an
ImportError — that error was from a stale local pipecat-ai package, not
a real drift. CI's pipecat build emits this frame and the test asserted
on it, so removing it broke test_llm_calls_custom_tool_handler and its
unregistered-call counterpart. Reverted both files to match main.
Also applied ruff's import-sort/format fix to test_mcp_tool_route.py to
clear the drift-check job (split the aliased import into its own
`from ... import (...)` block, wrapped a long monkeypatch.setattr call).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: avoid pytest collecting test_tool import as a test, sync OpenAPI spec
pytest's default discovery matches any top-level test_* name in a test
module, including imported functions — importing the route handler as
`test_tool as test_tool_route` still matched the pattern, so pytest
tried to run it as a test and failed injecting fixtures for tool_uuid/
request/user. Renamed the alias to call_test_tool_route.
Also regenerated docs/api-reference/openapi.json for the new
POST /tools/{tool_uuid}/test route and its two schemas (couldn't run
the dump script locally — pipecat-ai version mismatch documented
separately — so hand-built the diff to exactly match FastAPI's
get_openapi() output format, verified against neighboring routes).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address cubic review findings on test panel
- Resolve dotted context-variable keys into nested objects before
posting the test request. render_template's get_nested_value walks
nested dicts, so a flat key like "runtime_configuration.realtime_model"
never matched — templates referencing nested context always resolved
to empty.
- Restrict isHttpApiTool to an explicit category equality check instead
of inferring it from exclusions. native/integration tools are
currently disabled in the create-tool UI so this wasn't reachable
today, but the exclusion list silently goes stale as new categories
are added.
- Stop showing a green success badge for non-2xx responses.
execute_http_tool returns status: "success" for any HTTP exchange
that completes, regardless of status code — only transport-level
errors (timeout, connection failure) get status: "error". The test
panel now checks status_code is in the 2xx range before treating the
call as a success.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: address second round of cubic/greptile findings
- Guard setNestedValue against prototype-pollution keys (__proto__,
constructor, prototype) in the dotted context-var path before
traversing.
- Normalize status to "error" in the test route when the upstream
status_code is >= 400. execute_http_tool only distinguishes
transport-level failures (timeout, connection error) from
"success" — a completed 4xx/5xx exchange still came back as
"success" from the executor.
- Seed testArgValues defaults for number/boolean parameters via a
useEffect keyed on the parameters array, so a required number or
boolean field isn't silently omitted from the test request if the
user never touches its input.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: only seed test-arg defaults for required number/boolean params
Seeding optional number/boolean parameters silently changed the test
request — an optional boolean flag the tester never touched was sent
as true, which can flip upstream behavior unintentionally. Restrict
seeding to required parameters.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: match whitespace and fallback-filter syntax in context var detection
extractContextVars required an exact {{initial_context.foo}} with no
whitespace and no filter suffix, but the backend's TEMPLATE_VAR_PATTERN
(and render_template) accepts {{ initial_context.foo }} and
{{initial_context.foo | fallback:value}}. A preset parameter saved with
either of those forms resolved fine in production but showed no input
in the test panel, so testing always sent it empty context.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(tool-test): add hint and request_* fields to ToolTestResponse
Extends ToolTestResponse with hint, request_method, request_url,
request_body, and request_params so the frontend can surface what was
actually sent and a human-readable hint about why a test call failed.
* feat(tool-test): add status-code hints and request_method/url/body/params to test_tool()
* style: ruff-format test_mcp_tool_route.py
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(tool-test): wire hint banner and request block into result panel
Backend has returned hint/request_method/request_url/request_body/
request_params since d45ea851/60aaf31d but the frontend never
displayed them. Extends ToolTestResult with the new fields and renders
an amber hint banner (for 400/401/403/404/405/408/409/415/422/429/5xx)
plus a Request block above the response showing exactly what was sent.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(tool-test): include resolved preset params in request_body/params
Found via live GET/POST testing: the Request preview showed only the
model-provided arguments, not what execute_http_tool actually sends.
execute_http_tool merges resolved_arguments = {**arguments,
**preset_arguments} before building the outbound body/params — preset
params (e.g. {{initial_context.metadata.channel}}) are invisible to
the model but still go out on the wire. The preview now mirrors that
merge via the same _resolve_preset_parameters helper, so a dev sees
exactly what was sent, not just what the model provided.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(tool-test): add generateSampleValue helper for sample-fill button
* feat(tool-test): add Fill sample values button for arguments and context vars
Moved generateSampleValue out of page.tsx into a sibling helpers module:
Next.js's typed-route checker rejects extra named exports on a page.tsx
file (tsc error TS2344 on .next/types), so the helper and its test import
now live in testPanelHelpers.ts instead.
* feat(tool-test): add JSON edit modal state and handlers
* feat(tool-test): collapsed preview + edit modal for object/array test parameters
* fix(tool-test): validate JSON on modal open, not just on edit
Opening the JSON edit modal on an untouched object/array param (no value
yet in testArgValues) loaded an empty draft with jsonEditError hardcoded
to null, so Save was enabled despite invalid JSON and silently no-op'd on
click. Now runs the same JSON.parse check used by the live textarea
validation when the modal opens.
* chore: regenerate openapi.json for ToolTestResponse hint/request_* fields
drift-check on PR #547 was failing because the earlier hint/request_method/
request_url/request_body/request_params fields added to ToolTestResponse
were never reflected in the dumped spec. Regenerated via
scripts.dump_docs_openapi.
* fix(tool-test): serialize object/array args for GET/DELETE query params, add unsaved-changes banner
httpx raises a TypeError when a query param value is a dict/list, which
was silently caught and surfaced as a generic tool-execution error —
this is what actually broke test requests, not just the "[object
Object]" display. JSON-stringify object/array arguments before they
become query params, in both the live execute_http_tool() path and the
test route's request_params display shaping.
Also adds an unsaved-changes warning banner above Test Tool, shown
when the live form state diverges from the last-saved HTTP API config,
since Test Tool always runs the saved config.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(tool-test): keep empty request body in preview; normalize headers in snapshot
- POST/PUT/PATCH with no arguments now shows `{}` in the request body
preview instead of null — matches what execute_http_tool actually sends
over the wire (json={})
- buildHttpToolTestSnapshot normalizes headers from KeyValueItem[] to a
deduped key→value map before serializing, matching the shape saved to
the backend; duplicate header keys no longer cause a false unsaved-
changes warning
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* test(tool-test): update assertion for empty POST body preview
test_tool_test_no_arguments_leaves_body_and_params_none expected
request_body=None for a POST with no arguments. The fix to preserve {}
in the preview makes request_body={} the correct assertion.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat(tools): refine HTTP tool testing
---------
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>
* fix(web): honor X-Forwarded-Proto in uvicorn so request.url is https behind a reverse proxy
## Problem
When Dograh runs behind a TLS-terminating reverse proxy (Cloudflare →
Traefik in Kubernetes, nginx in the docker-compose install), the inside
of the cluster/host is plain HTTP. Uvicorn defaults to trusting
`scope["scheme"]` from the socket, so `request.url.scheme` reads `http`
even though the client dialed `https`.
That breaks any code path that hashes or echoes the request URL back to
the caller. Concrete symptom seen in production: **Vobiz inbound webhook
signatures fail with "signature validation failed for vobiz"** because
Vobiz computes HMAC over the URL it dialed (`https://.../inbound/run`)
while Dograh recomputes it as `http://...`. Log excerpt from the
failing call:
```
WARNING | provider.py | Vobiz webhook signature mismatch.
Expected: daOpAZPm..., Got: 1+eW/RxE...
WARNING | telephony.py | /inbound/run: signature validation failed for vobiz
```
Twilio, Plivo and any other provider that signs over the callback URL
have the same failure mode when Dograh is deployed behind a proxy.
## Fix
Start uvicorn with `--proxy-headers --forwarded-allow-ips="*"` in
`scripts/run_web.sh`. Uvicorn rewrites `scope["scheme"]` and client
address from `X-Forwarded-Proto` / `X-Forwarded-For` when the request
originates from a trusted upstream — Traefik and Cloudflare set both
correctly, so `request.url.scheme == "https"` inside the app once again
and provider signature checks pass.
Verified end-to-end on a production k3s install (Traefik + Cloudflare
edge → dograh-web pod) — after the change, the very next Vobiz inbound
webhook validated successfully and the call connected past the previous
11-second signature-failure hangup.
* address review: let operators narrow FORWARDED_ALLOW_IPS
Both bot reviewers on #515 flagged `--forwarded-allow-ips="*"` as a
defence-in-depth concern: if uvicorn is directly reachable from an
untrusted network (bypassing the proxy), any client can spoof
`X-Forwarded-Proto` / `X-Forwarded-For`, and uvicorn will rewrite
`request.client` / `request.url` from those attacker-controlled headers.
Fix: consume `FORWARDED_ALLOW_IPS` from the environment (uvicorn already
recognizes this env var; see `deploy/hostinger/docker-compose.yaml:179`
for the existing precedent). Default stays `"*"` so the behavior of the
original fix is preserved for the standard docker-compose / helm layouts
where the app pod is only reachable via the proxy Service. Operators
who terminate uvicorn on a host that's also reachable directly can
narrow it to the proxy CIDR:
FORWARDED_ALLOW_IPS="10.42.0.0/16" ./scripts/run_web.sh
* address review: declare FORWARDED_ALLOW_IPS in the helm chart, not the script
uvicorn already enables proxy-header handling by default and falls back to
the FORWARDED_ALLOW_IPS env var when --forwarded-allow-ips is absent, so the
CLI flags were redundant and the script-level "*" default hid a
security-relevant trust decision away from operators. Drop the flags, keep
run_web.sh deployment-agnostic, and declare the env var where the other
deployment config lives — web.forwardedAllowIps in values.yaml (default "*",
narrowable to a proxy CIDR) — mirroring how docker-compose already sets it
on the api service.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* simplify run_web.sh comment
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: prabhat pankaj <prabhatiitbhu@gmail.com>
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* docs: add video-embedded getting-started pages for API Trigger, Webhook, Telephony, Tools & Knowledge Base
Four new tutorial pages inserted after Your First Agent in 5 Minutes, each
pairing a walkthrough video with a step-by-step practical guide sourced
from the recorded demo: Trigger Calls Automatically (API Trigger), Send
Call Data Back Automatically (Webhook), Connect Your Phone Number
(Twilio telephony), and Give Your Agent Real Data (HTTP tools + KB).
* docs: address greptile review feedback on PR #535
* docs: link Twilio Verified Caller IDs page directly
* Restructure the documents
* docs: embed agent builder walkthrough video on first-agent page
* docs: match link text to renamed Connect with Telephony title
* docs: warn that default outbound telephony config is required for API Trigger
---------
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>
* Add support for foreground debugging
* Add support for Cloudonix call transfers
* Improve the customer/agent conference experience with less annoying sounds.
* Update remote_up.sh
Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com>
* Resolve a small redundant code segment from cubic
* Resolve an issue with callbacks not providing the correct experience for failed
originated calls
* Yet a small fix
* Remove stale code
* Remove the beeps on transfer
* Remove unrelated remote_up.sh changes
* Update pipecat submodule to main
---------
Co-authored-by: Nir Simionovich <nirs@cloudonix.com>
Co-authored-by: cubic-dev-ai[bot] <191113872+cubic-dev-ai[bot]@users.noreply.github.com>
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>
* fix(quota): fail closed when quota verification errors (#331)
Quota enforcement fell open on unexpected errors: the outer `except` in
`authorize_workflow_run_start` returned `has_quota=True`, so a degraded
database or a config-resolution bug let a billable run start unverified.
Billing and abuse protection are control-plane functions, so this is the
wrong default under exactly the degraded conditions that matter.
- Fail closed by default: the outer handler now returns
`has_quota=False` / `quota_check_failed`, reusing the existing message.
- Add `QUOTA_FAIL_MODE=closed|open` (default `closed`) so OSS self-hosters
can explicitly opt back into availability; the open path logs loudly.
- Narrow the try-scope so `get_user_by_id` / `get_workflow_run` DB read
failures surface as their specific `user_not_found` /
`workflow_run_not_found` codes instead of the generic handler.
- Tests cover the config-resolution and DB-read failure paths (denied,
not `has_quota=True`) and the `QUOTA_FAIL_MODE=open` escape hatch.
Fixes#331
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(quota): route DB read failures through the fail-mode policy gate
Review (greptile) flagged that the narrowed get_user_by_id / get_workflow_run
catches returned user_not_found / workflow_run_not_found before the outer
QUOTA_FAIL_MODE handler ran, so QUOTA_FAIL_MODE=open never applied to a DB
failure -- the exact "degraded database" case the escape hatch documents.
Revert the two narrowed catches so DB read exceptions fall through to the
single outer policy gate: closed -> quota_check_failed, open -> allow. The
None checks still return the specific not_found codes for genuinely missing
rows; an exception is a "cannot verify" condition, not a definitive absence.
Add a regression test asserting QUOTA_FAIL_MODE=open allows a run when a DB
read throws, and update the two DB-error tests to expect quota_check_failed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* docs(quota): scope the fail-mode comment to credit-verification failures (#331)
Review (cubic) flagged the outer-handler comment as overclaiming: it said the
handler is the single gate for "all cannot-verify errors", but the earlier
workflow-load and org-membership catches always deny with workflow_not_found
regardless of QUOTA_FAIL_MODE. That distinction is intentional (those are
authorization/existence gates, not credit verification), so scope the comment
accordingly. Comment-only, no behavior change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(quota): fail open only when MPS is unreachable
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Abhishek Kumar <abhishek@a6k.me>