Commit graph

2671 commits

Author SHA1 Message Date
Anish Sarkar
4fac1a20c3 feat(proxy): add geo country routing to proxy urls 2026-07-15 15:31:38 +05:30
Anish Sarkar
94b16a8334 fix(amazon): update anti-bot detection logic to include soft responses 2026-07-15 13:00:53 +05:30
Anish Sarkar
603c94c80b feat(amazon): enhance product fetching and parsing logic for better offer retrieval
- Updated the fetch function to correctly construct the URL for the all-offers panel.
- Improved the product parser to include additional selectors for price and seller information.
- Added a fallback mechanism to retrieve seller information from the buy-box when not available in the standard link.
2026-07-15 11:04:46 +05:30
Anish Sarkar
69848b4c71 feat(amazon): register amazon subagent 2026-07-15 10:24:45 +05:30
Anish Sarkar
6888f0b8bb feat(amazon): add subagent builder 2026-07-15 10:24:40 +05:30
Anish Sarkar
5a34435160 docs(amazon): add subagent description and system prompt 2026-07-15 10:24:34 +05:30
Anish Sarkar
03cd99194b feat(amazon): add subagent capability tools 2026-07-15 10:24:22 +05:30
Anish Sarkar
21a7a0a0b0 feat(amazon): register scrape capability 2026-07-15 01:57:32 +05:30
Anish Sarkar
3e0caa2027 feat(amazon): add scrape capability schema, executor, definition 2026-07-15 01:57:24 +05:30
Anish Sarkar
118411ebca docs(amazon): add module status readme 2026-07-15 01:56:34 +05:30
Anish Sarkar
f69a573d10 feat(amazon): expose public module api 2026-07-15 01:56:29 +05:30
Anish Sarkar
0cb15272ea feat(amazon): add scraper orchestrator flows 2026-07-15 01:56:19 +05:30
Anish Sarkar
165f7464b3 feat(amazon): add proxy-aware fetch with captcha rotate 2026-07-15 01:56:02 +05:30
Anish Sarkar
1b6383f9bc feat(amazon): add pure html parsers 2026-07-15 01:55:22 +05:30
Anish Sarkar
a0ab303768 feat(amazon): add start-url resolver 2026-07-15 01:55:15 +05:30
Anish Sarkar
4e5b039a05 feat(amazon): add scrape input and output schemas 2026-07-15 01:55:09 +05:30
Anish Sarkar
dafb0099c9 feat(proxy): add sticky-session proxy url support 2026-07-15 01:55:01 +05:30
Anish Sarkar
01fd6172ce feat(billing): add amazon product billing unit and rate 2026-07-15 01:54:49 +05:30
DESKTOP-RTLN3BA\$punk
2b018c4474 refactor: streamline TikTok and Instagram scraping logic by removing search_queries and enhancing documentation for clarity 2026-07-13 17:11:25 -07:00
DESKTOP-RTLN3BA\$punk
1131da5ed7 feat: bumped version to 0.0.32 2026-07-13 16:29:39 -07:00
Rohan Verma
d3da875658
Merge pull request #1599 from Thibaultjaigu/feat/requesty-provider
feat: add Requesty as a model provider
2026-07-13 13:03:20 -07:00
Rohan Verma
d1fa9306cb
Merge pull request #1597 from zzh-github/feature/openai-compatible-raw
Add raw OpenAI-compatible provider without /v1 normalization
2026-07-13 13:02:19 -07:00
Thibault Jaigu
2ff7ea4cb6 feat: add Requesty as a model provider
Add Requesty (https://requesty.ai), an OpenAI-compatible LLM router, as a
model provider by mirroring the existing OpenRouter integration.

Backend:
- app/services/requesty_model_normalizer.py: normalizes Requesty's /v1/models
  catalogue, mapping its flat capability booleans (supports_tool_calling/
  supports_vision/supports_image_generation) and context_window field onto the
  shared normalized shape (Requesty differs from OpenRouter's architecture +
  supported_parameters + context_length layout)
- provider_registry.py: Requesty ProviderSpec (OpenAI-compatible, base URL
  https://router.requesty.ai/v1, REQUESTY_API_KEY bearer auth)
- model_connection_service.py: key verification + live model discovery
- quality_score.py: Requesty score entry
- unit tests mirroring the OpenRouter normalizer coverage

Frontend:
- Requesty provider icon + registration, metadata entry, and base-url hint

Signed-off-by: Thibault Jaigu <thibault.jaigu@gmail.com>
2026-07-13 09:42:30 +01:00
zzh
a172e9c162 Add raw OpenAI-compatible provider without /v1 normalization 2026-07-13 00:08:30 +08:00
Anish Sarkar
e38ca19b18 Merge remote-tracking branch 'upstream/dev' into feat/instagram-scraper 2026-07-11 04:31:11 +05:30
Anish Sarkar
819486ac46 docs(instagram): document PolarisMedia extraction and richer media fields 2026-07-11 03:44:38 +05:30
Anish Sarkar
d3c65a37b1 refactor(instagram): drop unused slug field and mark share links unsupported 2026-07-11 03:44:28 +05:30
Anish Sarkar
4bc3d43b56 refactor(instagram): remove unused share-link redirect resolver 2026-07-11 03:44:23 +05:30
Anish Sarkar
452fe30aa4 refactor(instagram): drop login-walled comment fields and note search_type aliasing 2026-07-11 03:44:15 +05:30
Anish Sarkar
cb4adab852 feat(instagram): extract PolarisMedia relay JSON with carousel, tags, and location 2026-07-11 03:44:10 +05:30
Anish Sarkar
b1fe739308 docs(instagram): update platform scraper README 2026-07-11 03:25:36 +05:30
Anish Sarkar
c001f4b16e feat(subagent): update Instagram subagent system prompt 2026-07-11 03:25:25 +05:30
Anish Sarkar
27d22a9a2a refactor(instagram): drop mentions from scrape capability 2026-07-11 03:25:20 +05:30
Anish Sarkar
5cac6612c3 refactor(instagram): update platform schemas and scraper for tagged media 2026-07-11 03:25:17 +05:30
Anish Sarkar
457be1871c feat(instagram): improve parsing of Instagram media IDs and mentions
- Enhanced the regular expression for mentions to prevent trailing punctuation from being included in handles.
- Added support for extracting media IDs from deep-link meta tags in anonymous posts.
- Updated unit tests to validate the new media ID extraction and ensure proper handling of mentions.
2026-07-11 03:06:04 +05:30
Anish Sarkar
d41ccb7edd feat(instagram): enhance Open Graph parsing for anonymous posts
- Added support for extracting likes, comments, username, timestamp, and caption from Open Graph meta tags.
- Implemented fallback mechanisms to ensure graceful degradation when expected data is missing.
- Updated unit tests to validate new parsing logic and ensure robustness against unrecognized formats.
2026-07-11 02:56:29 +05:30
Anish Sarkar
9e9dd8a124 feat(subagent): update Instagram subagent prompts and tools 2026-07-11 02:21:02 +05:30
Anish Sarkar
36bb4a1bea docs(instagram): update platform scraper README 2026-07-11 02:20:54 +05:30
Anish Sarkar
4813dc96e3 refactor(instagram): streamline scraper, fetch, and parsing pipeline 2026-07-11 02:20:51 +05:30
Anish Sarkar
de1990f9f6 refactor(instagram): simplify scrape and details capability schemas 2026-07-11 02:20:46 +05:30
Anish Sarkar
8afa4c6fc6 refactor(instagram): remove unused comments capability 2026-07-11 02:20:39 +05:30
CREDO23
1911b78379 fix(tiktok-agent): drop keyword search from the routing description's video sources 2026-07-10 17:40:04 +02:00
CREDO23
cd82f861de fix(tiktok-agent): stop steering the subagent to keyword search for videos
The playbook listed search_queries as a video source in three places, then contradicted itself. Replace with action-oriented guidance: videos via hashtags/URLs, accounts via user_search. Drop 'login-walled' (the agent has no login path).
2026-07-10 17:40:04 +02:00
CREDO23
f74a73efdb docs(tiktok): trim fetch_item_list retry docstring to intent 2026-07-10 17:28:47 +02:00
CREDO23
01bc3f8de3 docs(config): trim narration from xvfb + listing-retry comments 2026-07-10 17:28:47 +02:00
CREDO23
c8fcf87685 fix(tiktok): retry empty item_list captures on a fresh exit IP
The profile feed is withheld from flagged IPs and the proxy rotates per request, so re-fetching an empty capture (up to TIKTOK_LISTING_MAX_ATTEMPTS) turns a bad first draw into a hit instead of an ErrorItem.
2026-07-10 17:25:48 +02:00
CREDO23
15fb50c83e feat(tiktok): add TIKTOK_LISTING_MAX_ATTEMPTS knob
Bounds the browser-listing retry-on-empty; default 3, set 1 for a static IP.
2026-07-10 17:25:48 +02:00
CREDO23
6400dc5f04 fix(tiktok): gate headful profile feed on CRAWL_HEADED_XVFB_ENABLED
fetch_item_list goes headful only when the flag promises a display; otherwise it stays headless so the browser launch never fails and the empty feed degrades to an ErrorItem. Verified under xvfb-run: real videos returned on a clean proxy IP, graceful degrade on a flagged one.
2026-07-10 17:17:45 +02:00
CREDO23
6784503414 feat(crawl): add CRAWL_HEADED_XVFB_ENABLED flag
Single switch promising an Xvfb display so the stealth browser can run headful (TikTok's profile feed is empty to headless Chromium). Defaults FALSE.
2026-07-10 17:17:27 +02:00
CREDO23
c679f2a3ef fix(tiktok): fetch profile video lists headful + dismiss login modal
The profile feed (/api/post/item_list) returns an empty 200 to headless
sessions but serves data headful on the same proxy IP. Run fetch_item_list
headful and dismiss the login modal that blocks mid-scroll.
2026-07-10 16:45:45 +02:00