Commit graph

76 commits

Author SHA1 Message Date
DESKTOP-RTLN3BA\$punk
dbedf0cfa5 feat(backend): integrate PostHog analytics for enhanced observability
- Added PostHog configuration options to .env.example files for both Docker and Surfsense backend.
- Introduced PostHog dependency in pyproject.toml.
- Implemented analytics middleware to capture various events across the application, including user authentication, automation runs, and API requests.
- Enhanced existing routes and services to emit analytics events, providing insights into user interactions and system performance.
- Ensured graceful shutdown of analytics clients in worker processes and application lifecycles.
2026-07-22 22:16:28 -07:00
CREDO23
950de5e670 fix: forward ToolRuntime to capability tools 2026-07-22 20:06:18 +02:00
CREDO23
c1ba4215fd feat: mint run citation from capability tool 2026-07-22 19:34:56 +02:00
CREDO23
957e57bb82 feat: register a scraper run as a citation 2026-07-22 19:34:56 +02:00
CREDO23
ca317c4686 Merge remote-tracking branch 'upstream/dev' into feature-walmart-scraper
# Conflicts:
#	docker/.env.example
#	surfsense_backend/app/capabilities/core/billing.py
#	surfsense_backend/app/capabilities/core/types.py
#	surfsense_backend/app/config/__init__.py
#	surfsense_mcp/mcp_server/features/scrapers/__init__.py
#	surfsense_web/content/docs/connectors/index.mdx
#	surfsense_web/content/docs/connectors/native/index.mdx
#	surfsense_web/content/docs/how-to/mcp-server.mdx
2026-07-21 02:42:41 +02:00
CREDO23
d7c30cac2f refactor(walmart): align reviews executor docstring 2026-07-19 08:45:27 +02:00
CREDO23
9b9091ae97 feat(walmart): add scrape and reviews capabilities 2026-07-19 08:19:22 +02:00
CREDO23
0d2c8df5bd feat(walmart): add billing units and rate config 2026-07-19 08:19:21 +02:00
CREDO23
91aa265afb Merge remote-tracking branch 'upstream/dev' into feature-indeed-jobs-scraper
# Conflicts:
#	README.es.md
#	README.hi.md
#	README.md
#	README.pt-BR.md
#	README.zh-CN.md
#	surfsense_backend/tests/unit/capabilities/google_maps/test_registry.py
#	surfsense_backend/tests/unit/capabilities/reddit/test_registry.py
#	surfsense_backend/tests/unit/capabilities/youtube/test_registry.py
#	surfsense_mcp/mcp_server/features/scrapers/__init__.py
#	surfsense_web/content/docs/connectors/index.mdx
#	surfsense_web/content/docs/connectors/native/index.mdx
#	surfsense_web/content/docs/connectors/native/meta.json
#	surfsense_web/content/docs/how-to/mcp-server.mdx
#	surfsense_web/lib/connectors-marketing/index.ts
#	surfsense_web/lib/playground/catalog.ts
2026-07-18 22:11:30 +02:00
CREDO23
ec258d8387 feat: map scrape_job_details in indeed executor 2026-07-15 17:54:38 +02:00
CREDO23
81f535e9d0 feat: add scrape_job_details to indeed.scrape input 2026-07-15 17:54:38 +02:00
Anish Sarkar
21a7a0a0b0 feat(amazon): register scrape capability 2026-07-15 01:57:32 +05:30
Anish Sarkar
3e0caa2027 feat(amazon): add scrape capability schema, executor, definition 2026-07-15 01:57:24 +05:30
Anish Sarkar
01fd6172ce feat(billing): add amazon product billing unit and rate 2026-07-15 01:54:49 +05:30
CREDO23
8c410d8682 feat: indeed capability namespace 2026-07-14 21:49:23 +02:00
CREDO23
17159172ee feat: indeed.scrape verb package 2026-07-14 21:49:23 +02:00
CREDO23
3f91a006cd feat: register indeed.scrape capability 2026-07-14 21:49:23 +02:00
CREDO23
b457599fd4 feat: indeed.scrape executor 2026-07-14 21:49:23 +02:00
CREDO23
e887c1cae7 feat: indeed.scrape io schemas 2026-07-14 21:49:23 +02:00
CREDO23
d89ca141a4 feat: meter indeed_job billing unit 2026-07-14 21:49:23 +02:00
CREDO23
05bfc6a10e feat: add INDEED_JOB billing unit 2026-07-14 21:49:23 +02:00
DESKTOP-RTLN3BA\$punk
2b018c4474 refactor: streamline TikTok and Instagram scraping logic by removing search_queries and enhancing documentation for clarity 2026-07-13 17:11:25 -07:00
DESKTOP-RTLN3BA\$punk
1131da5ed7 feat: bumped version to 0.0.32 2026-07-13 16:29:39 -07:00
Anish Sarkar
e38ca19b18 Merge remote-tracking branch 'upstream/dev' into feat/instagram-scraper 2026-07-11 04:31:11 +05:30
Anish Sarkar
27d22a9a2a refactor(instagram): drop mentions from scrape capability 2026-07-11 03:25:20 +05:30
Anish Sarkar
de1990f9f6 refactor(instagram): simplify scrape and details capability schemas 2026-07-11 02:20:46 +05:30
Anish Sarkar
8afa4c6fc6 refactor(instagram): remove unused comments capability 2026-07-11 02:20:39 +05:30
CREDO23
67b5472b9f feat(tiktok): add tiktok.trending verb for the Explore feed
The Explore feed (/api/explore/item_list) is a global trending-video feed
served to anonymous sessions, and it returns the same itemStruct shape as the
other listings — so the verb reuses parse_video, the listing flow, the
TikTokVideoItem output, and the per-video billing meter wholesale. Adds a
browser-capture marker + fetch_trending, a synthetic-target orchestrator entry,
and the tiktok.trending capability, surfaced on the chat subagent.
2026-07-09 18:52:56 +02:00
CREDO23
7723b5b8b6 feat(tiktok): add tiktok.comments verb for public comment scraping
Comments load over a signed /api/comment/list XHR that TikTok serves to
anonymous sessions once the comments panel is opened (unlike profile-video and
general-search feeds), so this is a reliable verb. Given video URLs it returns
CommentItems (text, author, likes, reply counts; replies carry repliesToId),
deduped per video, capped, and degraded to an ErrorItem for empty/withheld
videos or a bad_url ErrorItem for non-video inputs.

Generalizes the browser capture over a pluggable interaction step so the
comments flow (open panel, scroll the panel to paginate) reuses the same
warm+capture scaffolding as listing/user-search. Billed per comment on a new
TIKTOK_COMMENT meter (TIKTOK_MICROS_PER_COMMENT, matching the per-comment
market), surfaced on the chat subagent alongside tiktok.scrape/user_search.
2026-07-09 18:35:26 +02:00
CREDO23
192b6dc31a feat(tiktok): add tiktok.user_search verb for account discovery
Video/general search is login-walled for anonymous sessions, but the Users
tab (/api/search/user) returns public account records without a redirect, so
this exposes the one reliably-unblocked search path. A keyword yields
TikTokProfileItems (name, followers, bio, verification), deduped per query,
capped, and degraded to an ErrorItem when a query is empty/withheld.

Reuses the browser capture (generalized over XHR markers + extractor) and the
shared profile item shape. Billed per account on a new TIKTOK_USER meter
(TIKTOK_MICROS_PER_USER), surfaced on the chat subagent alongside tiktok.scrape.
2026-07-09 18:00:40 +02:00
Anish Sarkar
82f1d0b4e5 refactor(instagram): improve anonymous scraping logic and error messaging
Enhanced the Instagram scraper to clarify the requirements for accessing user profiles and hashtags. Updated the error message for blocked access to provide detailed guidance on necessary credentials. Introduced a regex for validating Instagram usernames and refined the discovery function to handle profile queries directly, improving user experience and error handling in anonymous mode.
2026-07-09 20:23:09 +05:30
Anish Sarkar
929efe8152 feat(instagram): register instagram.* capability namespace 2026-07-09 16:01:49 +05:30
Anish Sarkar
7ea474be93 feat(instagram): add details capability verb 2026-07-09 16:01:38 +05:30
Anish Sarkar
38a26b1ccf feat(instagram): add comments capability verb 2026-07-09 14:58:58 +05:30
Anish Sarkar
ff55537bce feat(instagram): add scrape capability verb 2026-07-09 14:58:54 +05:30
Anish Sarkar
43660ee51b feat(billing): map Instagram units to rate keys and display nouns 2026-07-09 14:57:53 +05:30
Anish Sarkar
f204db7bce feat(billing): add INSTAGRAM_ITEM and INSTAGRAM_COMMENT units 2026-07-09 14:57:44 +05:30
CREDO23
ba375ea3ee feat(tiktok): emit graceful ErrorItem for blocked/empty listings
Profile and search feeds are trust-gated: an anonymous headless session
gets an empty item_list (profile) or no results XHR (search), while
hashtag feeds load. A zero-item listing now yields one honest ErrorItem
(errorCode="no_items") instead of vanishing silently, and ErrorItems are
excluded from billing so a blocked target is surfaced but never charged.
2026-07-08 23:14:50 +02:00
CREDO23
2943d8b23c feat(tiktok): tiktok.scrape capability + billing wire-up 2026-07-08 18:21:44 +02:00
Anish Sarkar
7ba77c6b86 feat(capabilities): add documentation URLs to Google Maps, YouTube, and web crawl capabilities
- Introduced a `docs_url` field to the Google Maps reviews, scrape, YouTube comments, YouTube scrape, and web crawl capabilities for improved documentation access.
- Simplified descriptions for each capability to enhance clarity and user understanding.
2026-07-08 03:38:08 +05:30
Anish Sarkar
3b9c806c90 feat(capabilities): add documentation URL to capabilities and enhance Reddit scrape capability
- Introduced a `docs_url` field to the `Capability` and `CapabilitySummary` classes for better documentation access.
- Updated the `REDDIT_SCRAPE` capability to include a specific documentation link.
- Enhanced the PlaygroundRunner component to display the documentation link when available, improving user guidance.
2026-07-08 03:35:32 +05:30
DESKTOP-RTLN3BA\$punk
1c9ab207ef chore: bumped version to 0.0.31 2026-07-06 21:43:15 -07:00
DESKTOP-RTLN3BA\$punk
271a21aee6 feat: docs and ui tweaks 2026-07-06 19:26:35 -07:00
DESKTOP-RTLN3BA\$punk
50f2d095aa feat: api playground, other tweaks 2026-07-06 16:45:04 -07:00
DESKTOP-RTLN3BA\$punk
80927a2872 feat(billing): meter platform scrapers per item; consolidate web scraping onto web.crawl
Add per-item, per-platform billing for the platform-native connectors (Reddit, Google Search, Google Maps places/reviews, YouTube videos/comments) through the capability gate/charge seam. Rates are config-driven with a shared wallet-credit module (wallet_credit) and a dedicated PlatformScrapeCreditService; agent and REST capability runs now record cost_micros. Google Maps scrape dual-meters places and attached reviews.

Remove the main-agent scrape_webpage tool now that the web.crawl capability covers single-page (maxCrawlDepth=0) and site crawling. The main agent now reaches crawling via task(web_crawler, ...). Update prompts, tool catalog, receipts, skills, proprietary docs, and tests; drop the obsolete chat-turn crawl fold path.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-05 17:08:01 -07:00
DESKTOP-RTLN3BA\$punk
c600a2920b feat(crawler): harden web crawler and agent tooling for real-world research tasks
Crawler engine: escalate thin JS-shell pages past static fetch, repair
currency-lossy extractions, emit categorized link records with anchor
provenance, and decode percent-encoded mailto:/tel: contacts; site crawls
reuse the connector ladder via Scrapling's spider engine with URL pattern
filters. Agent layer: read_run gains char_offset paging, search_run gains
match excerpts, new export_run turns stored runs into CSV workspace docs;
reddit search fair-shares the item budget across queries and dedupes
cross-query hits. Subagent prompts and routing teach crawl-after-search,
full-run coverage before summarizing, and executing own-tool next steps
instead of returning partial.
2026-07-05 03:51:16 -07:00
DESKTOP-RTLN3BA\$punk
b6e378b070 feat(runs): introduce Run and ToolOutputSpill models for logging scraper invocations
- Added `Run` model to track scraper invocations, including metadata such as status, input, and output.
- Implemented `ToolOutputSpill` model to store context-editing spills separately from user-facing logs.
- Updated middleware to handle spill placeholders and integrate with the new models.
- Enhanced REST API to record runs and expose run history through new endpoints.
- Adjusted tests to validate the new run logging functionality and ensure proper integration with existing capabilities.
2026-07-04 22:55:55 -07:00
DESKTOP-RTLN3BA\$punk
ff2e5f390f feat(reddit): implement Reddit scraping subagent and associated capabilities
- Added a new `reddit` subagent to scrape structured data from Reddit posts, comments, and users.
- Introduced `reddit.scrape` capability for fetching data using URLs and search queries.
- Implemented tools for scraping and parsing Reddit data, including handling pagination and rate limits.
- Created input/output models for the Reddit scraper to define request and response structures.
- Added documentation for the new Reddit scraping functionality and its usage.
- Integrated the Reddit subagent into the existing multi-agent chat framework.
2026-07-04 17:31:11 -07:00
CREDO23
b4ad8c5f4c feat(google_search): register scrape capability 2026-07-04 20:49:58 +02:00
CREDO23
fcf1c6e052 feat(google_search): add scrape executor 2026-07-04 20:49:58 +02:00