Rohan Verma
07ace7ff7d
Merge pull request #1571 from AnishSarkar22/feat/reddit-scraping
...
feat(scrapers): add anonymous Reddit scraper
2026-07-04 15:37:53 -07:00
Anish Sarkar
82141d8047
chore: fix docs
2026-07-04 17:45:43 +05:30
Anish Sarkar
0a83bf6192
test(native-connector): cover reddit schemas and url resolver
2026-07-04 17:28:21 +05:30
Anish Sarkar
0ef02c43cc
test(native-connector): cover reddit parser mappings
2026-07-04 17:28:17 +05:30
Anish Sarkar
1b1e612e20
test(native-connector): cover reddit fetch resilience
2026-07-04 17:28:13 +05:30
Anish Sarkar
4d950ddcb5
test(native-connector): add reddit post fixture
2026-07-04 17:28:08 +05:30
Anish Sarkar
672f79d620
test(native-connector): add reddit listing fixture
2026-07-04 17:28:02 +05:30
Anish Sarkar
83f52b17b1
test(native-connector): add reddit comment fixture
2026-07-04 17:27:56 +05:30
Anish Sarkar
f14bf1f06d
test(native-connector): add reddit scraper test package
2026-07-04 17:27:53 +05:30
pelazas
f44cb8a493
fix: dispatch periodic BookStack indexing from the scheduler task_map
...
BOOKSTACK_CONNECTOR was missing from check_periodic_schedules' task_map,
so connectors with periodic indexing enabled fell through to the
"No task found" warning branch every minute and never synced.
Fixes #891
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-04 09:53:51 +02:00
Rohan Verma
aa7388f2f7
Merge pull request #1563 from mvanhorn/fix/1531-video-presentation-fanin-barrier
...
fix: wire video_presentation fan-in as a barrier so generate_slide_scene_codes fires once
2026-07-03 22:26:32 -07:00
DESKTOP-RTLN3BA\$punk
0d83916cc5
feat(native-connector): added google search results scraper
2026-07-03 20:51:45 -07:00
CREDO23
b5d221c19c
fix(capabilities): map any executor fault to 502 at the REST door
2026-07-03 17:57:07 +02:00
CREDO23
2c76ef1f57
feat(subagents): add google_maps builtin subagent
2026-07-03 17:56:53 +02:00
CREDO23
f31823f765
feat(google_maps): add scrape + reviews capability verbs
2026-07-03 17:56:43 +02:00
CREDO23
d5d673384f
Merge upstream/ci_mvp (google maps scrapers) into feature-ci_phase-4-7
...
Resolve conflicts against the new native google-maps actor + repo-wide
ruff-format pass:
- Keep legacy webcrawler KB indexer + its test deleted (modify/delete).
- test_validators: keep WEBCRAWLER case removed (validator gone).
- test_fetch_resilience: keep platforms.youtube import path (our reorg).
- Relocate google_maps actor + tests scrapers/ -> platforms/ to match the
reorg convention (youtube already there); rewrite imports + fixture paths.
- Add missing __init__.py across the capabilities/ test subtree so duplicate
test basenames get unique module paths under importlib mode.
Note: google_maps fixture-backed tests error on ci_mvp too (fixtures/*.json
never committed upstream) - pre-existing, out of scope here.
2026-07-03 12:37:12 +02:00
CREDO23
817ec0e9f4
refactor(youtube): remove legacy KB ingestion; use actor get_youtube_video_id
2026-07-03 12:00:22 +02:00
CREDO23
ab6be6cbda
refactor(connectors): remove legacy web-crawler KB indexer paths (keep enums)
2026-07-03 11:53:33 +02:00
CREDO23
e3ed3b2be3
refactor(subagents): remove dormant research subagent
2026-07-03 11:43:26 +02:00
CREDO23
c66b2f0e0e
refactor(automations): remove phase-6 chat watch (watch service/routes + chat_message action)
2026-07-03 11:20:22 +02:00
CREDO23
5da624399d
refactor(subagents): split scraping into web_crawler + youtube builtins
2026-07-03 11:20:09 +02:00
CREDO23
3312a442f3
refactor(proprietary): move scrapers/ -> platforms/ (one subpackage per platform)
2026-07-03 11:01:58 +02:00
CREDO23
9fe9c5b71d
feat(capabilities): replace web.scrape/web.discover with unified web.crawl
...
web.crawl scrapes a single URL (maxCrawlDepth=0) or spiders a whole site,
backed by the proprietary site_crawler engine. Rewires the scraping subagent
tools and capability tests onto the new verb.
2026-07-03 10:51:05 +02:00
CREDO23
f82fae3973
feat(web-crawler): add site_crawler spider + url_policy over the fetch engine
2026-07-03 10:50:41 +02:00
DESKTOP-RTLN3BA\$punk
7185079bd6
feat(native-connector): added google maps places & reviews scrapers
2026-07-02 21:58:24 -07:00
CREDO23
5d0bd8125b
test(capabilities): cover youtube verb schemas, executors, and registry
2026-07-02 20:18:53 +02:00
CREDO23
af3e70ea56
Merge upstream/ci_mvp: add YouTube scraper actor under app/proprietary
...
Resolved modify/delete conflicts on the old 'revamp phases 4-7/' plan docs
by keeping our deletion (superseded by flattened plans/backend/00-*.md).
2026-07-02 20:01:11 +02:00
CREDO23
6bd0d640e6
refactor(capabilities): group framework into core/, keep verbs top-level
2026-07-02 19:54:20 +02:00
CREDO23
87df0bb629
fix(web.discover): cap results to top_k
2026-07-02 19:40:37 +02:00
CREDO23
9738392161
refactor(agents): rename intelligence_agent subagent to scraping
2026-07-02 18:59:34 +02:00
CREDO23
790507d107
feat(capabilities): meter captcha, truncate scrape, bound discover
2026-07-02 18:32:42 +02:00
CREDO23
18e2ad13ca
test(automations): cover checkpointer cross-loop durability
2026-07-02 17:54:31 +02:00
CREDO23
f0c8d9ba5c
test(chat): cover watch tools
2026-07-02 17:54:31 +02:00
CREDO23
f1ea60365c
test(automations): cover watch rest routes
2026-07-02 17:54:31 +02:00
CREDO23
555dc3fcad
test(automations): cover chat watch service
2026-07-02 17:54:31 +02:00
CREDO23
61412163fd
test(automations): cover chat_message action
2026-07-02 17:54:31 +02:00
CREDO23
fecb4fa05e
test(automations): add integration harness
2026-07-02 17:54:31 +02:00
CREDO23
3278f46585
fix(tasks): dispose checkpointer pool per celery task
2026-07-02 17:54:31 +02:00
DESKTOP-RTLN3BA\$punk
0445c646c6
refactor(native-connector): relocate youtube scraper under app/proprietary + parallelize playlists
...
Move app/scrapers -> app/proprietary/scrapers/youtube to sit alongside the existing proprietary web_crawler/platforms namespace, updating all external imports (routes, tests, e2e script, README). Internal imports were relative so are unchanged.
Also parallelize playlist per-video resolution: page video ids sequentially, then resolve the heavy watch-page fetches concurrently via fan_out (~150 videos ~70s, down from a few minutes). Items stream in completion order; sort by the order field for playlist order.
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-02 02:34:04 -07:00
DESKTOP-RTLN3BA\$punk
99378d7b63
feat(native-connector): added youtube scrapers
2026-07-02 01:10:31 -07:00
CREDO23
7ad83760b5
feat(chat): add intelligence_agent subagent over capability verbs
2026-07-02 00:17:22 +02:00
CREDO23
110beb4dd1
feat(chat): add web_scrape/web_discover tool emission handlers
2026-07-02 00:17:10 +02:00
CREDO23
de4b8ab6b0
feat(capabilities): generate agent tool door from registry
2026-07-02 00:17:01 +02:00
CREDO23
23802a74bb
feat(capabilities): document verbs with descriptions and field docs
2026-07-02 00:16:55 +02:00
CREDO23
b1bd35c082
test(capabilities): cover REST door, meter-gate, and batch bounds
2026-07-01 17:53:42 +02:00
CREDO23
f257e385db
feat(capabilities): register web.discover verb
2026-07-01 17:08:58 +02:00
CREDO23
329bc63b6c
feat(capabilities): add env-keyed searxng/linkup/baidu discover providers
2026-07-01 17:08:58 +02:00
CREDO23
fa1055dd4c
feat(capabilities): add web.discover schemas and provider seam
2026-07-01 17:08:58 +02:00
CREDO23
a413539f6a
feat(capabilities): allow free verbs with no billing unit
2026-07-01 17:08:58 +02:00
CREDO23
8d76afc9d6
test(capabilities): cover registry, executor, schemas, billing
2026-07-01 16:42:47 +02:00