feat(billing): meter platform scrapers per item; consolidate web scraping onto web.crawl

Add per-item, per-platform billing for the platform-native connectors (Reddit, Google Search, Google Maps places/reviews, YouTube videos/comments) through the capability gate/charge seam. Rates are config-driven with a shared wallet-credit module (wallet_credit) and a dedicated PlatformScrapeCreditService; agent and REST capability runs now record cost_micros. Google Maps scrape dual-meters places and attached reviews.

Remove the main-agent scrape_webpage tool now that the web.crawl capability covers single-page (maxCrawlDepth=0) and site crawling. The main agent now reaches crawling via task(web_crawler, ...). Update prompts, tool catalog, receipts, skills, proprietary docs, and tests; drop the obsolete chat-turn crawl fold path.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
DESKTOP-RTLN3BA\$punk 2026-07-05 17:08:01 -07:00
parent b8285a0b72
commit 80927a2872
48 changed files with 724 additions and 766 deletions

View file

@ -31,7 +31,7 @@ Apache-2.0* — instead of per-file headers scattered across the tree.
- **Do not** add Apache-2.0-intended code here.
- Apache-2.0 code elsewhere **may import from** this package (the indexer and the
chat `scrape_webpage` tools do); that does not move them under this license.
`web.crawl` capability do); that does not move them under this license.
- Depend only on the public API exported from each subpackage's `__init__`, not
on internal modules, so the boundary stays clean and swappable.
- **Boundary test:** put code here only if it is used *exclusively* by the moat.

View file

@ -77,8 +77,9 @@ provenance (depth, referrer).
## Agent tooling layer (outside this package)
- The main chat agent has `scrape_webpage`; the `web_crawler` subagent has the
`web.crawl` capability (single URL or site mode).
- The `web_crawler` subagent exposes the `web.crawl` capability (single URL at
`maxCrawlDepth=0`, or site mode at higher depth); the main chat agent reaches
it by delegating via `task(web_crawler, …)`.
- Tool outputs over the 40k-char cap (`RUN_OUTPUT_CHAR_CAP` in
`app/capabilities/core/runs.py`) are stored as JSONL runs; agents page them
with `read_run` (line paging + `char_offset` for giant single items), grep