mirror of
https://github.com/MODSetter/SurfSense.git
synced 2026-07-22 23:31:12 +02:00
Merge pull request #1605 from CREDO23/feature-indeed-jobs-scraper
[Feat] Add Indeed jobs scraper
This commit is contained in:
commit
351696bd2c
65 changed files with 4326 additions and 17 deletions
|
|
@ -22,7 +22,7 @@
|
|||
|
||||
# SurfSense: La alternativa de código abierto a NotebookLM para la investigación de la web abierta
|
||||
|
||||
SurfSense es la **alternativa de código abierto a NotebookLM para agentes de IA**, una plataforma de investigación de la web abierta con conectores de datos en vivo. Tus agentes investigan la web en vivo con datos estructurados de **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search y cualquier página de la web abierta**, a través de una única **API REST** o un **servidor MCP**. Agentes programados y activados por eventos convierten lo que encuentran en informes y alertas, y una base de conocimiento integrada mantiene cada hallazgo disponible para búsqueda con citas.
|
||||
SurfSense es la **alternativa de código abierto a NotebookLM para agentes de IA**, una plataforma de investigación de la web abierta con conectores de datos en vivo. Tus agentes investigan la web en vivo con datos estructurados de **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed y cualquier página de la web abierta**, a través de una única **API REST** o un **servidor MCP**. Agentes programados y activados por eventos convierten lo que encuentran en informes y alertas, y una base de conocimiento integrada mantiene cada hallazgo disponible para búsqueda con citas.
|
||||
|
||||
> [!NOTE]
|
||||
> **📢 Una nota para nuestros usuarios de la alternativa a NotebookLM**
|
||||
|
|
@ -60,6 +60,7 @@ Pregúntale a cualquier agente capaz "¿qué está diciendo Reddit sobre este pr
|
|||
| **TikTok** | Videos, comentarios, hashtags y perfiles sin aprobación de la Research API | [TikTok Scraper API](https://www.surfsense.com/tiktok) |
|
||||
| **Google Maps** | Lugares, calificaciones y reseñas para investigar negocios locales | [Google Maps Scraper API](https://www.surfsense.com/google-maps) |
|
||||
| **Google Search** | SERPs en vivo para investigación y monitoreo de búsquedas | [Google Search API](https://www.surfsense.com/google-search) |
|
||||
| **Indeed** | Ofertas de empleo públicas con salarios y descripciones completas, por búsqueda o empresa | [Indeed Scraper API](https://www.surfsense.com/indeed) |
|
||||
| **Amazon** | Datos públicos de productos: precios, calificaciones, ofertas, vendedores y rankings de más vendidos | [Amazon Product API](https://www.surfsense.com/amazon) |
|
||||
| **Web Crawl** (rastreo web) | Cualquier página de la web abierta como contenido limpio y estructurado | [Web Crawling API](https://www.surfsense.com/web-crawl) |
|
||||
| **Conectores MCP externos** | Conecta cualquier servidor MCP a tus agentes, con OAuth de un clic para Notion, Slack, Jira y más | [External MCP Connectors](https://www.surfsense.com/external-mcp-connectors) |
|
||||
|
|
@ -217,7 +218,7 @@ SurfSense es el único producto de código abierto que combina un espacio de tra
|
|||
|
||||
| Característica | Google NotebookLM | SurfSense |
|
||||
|---------|-------------------|-----------|
|
||||
| **Datos web en vivo para agentes** | No | Conectores de Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search y rastreo web vía API REST y MCP |
|
||||
| **Datos web en vivo para agentes** | No | Conectores de Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed y rastreo web vía API REST y MCP |
|
||||
| **Servidor MCP** | No | Cada conector expuesto como herramienta nativa de agente, más servidores MCP propios con aplicaciones OAuth de un clic |
|
||||
| **Fuentes por Notebook** | 50 (gratis) a 600 (Ultra, $249.99/mes) | Ilimitadas |
|
||||
| **Número de Notebooks** | 100 (gratis) a 500 (niveles de pago) | Ilimitado |
|
||||
|
|
|
|||
|
|
@ -22,7 +22,7 @@
|
|||
|
||||
# SurfSense: ओपन वेब रिसर्च के लिए ओपन सोर्स NotebookLM विकल्प
|
||||
|
||||
SurfSense **AI एजेंट्स के लिए ओपन सोर्स NotebookLM विकल्प** है, लाइव डेटा कनेक्टर्स के साथ एक ओपन वेब रिसर्च प्लेटफ़ॉर्म। आपके एजेंट **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search और ओपन वेब के किसी भी पेज** से स्ट्रक्चर्ड डेटा के साथ लाइव वेब पर रिसर्च करते हैं, वह भी एक ही **REST API** या **MCP सर्वर** के ज़रिए। शेड्यूल्ड और इवेंट-ट्रिगर्ड एजेंट अपनी खोजों को ब्रीफ़ और अलर्ट में बदलते हैं, और एक बिल्ट-इन नॉलेज बेस हर खोज को साइटेशन के साथ खोजने योग्य बनाए रखता है।
|
||||
SurfSense **AI एजेंट्स के लिए ओपन सोर्स NotebookLM विकल्प** है, लाइव डेटा कनेक्टर्स के साथ एक ओपन वेब रिसर्च प्लेटफ़ॉर्म। आपके एजेंट **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed और ओपन वेब के किसी भी पेज** से स्ट्रक्चर्ड डेटा के साथ लाइव वेब पर रिसर्च करते हैं, वह भी एक ही **REST API** या **MCP सर्वर** के ज़रिए। शेड्यूल्ड और इवेंट-ट्रिगर्ड एजेंट अपनी खोजों को ब्रीफ़ और अलर्ट में बदलते हैं, और एक बिल्ट-इन नॉलेज बेस हर खोज को साइटेशन के साथ खोजने योग्य बनाए रखता है।
|
||||
|
||||
> [!NOTE]
|
||||
> **📢 हमारे NotebookLM-विकल्प उपयोगकर्ताओं के लिए एक सूचना**
|
||||
|
|
@ -60,6 +60,7 @@ SurfSense **AI एजेंट्स के लिए ओपन सोर्स
|
|||
| **TikTok** | Research API अप्रूवल के बिना वीडियो, कमेंट, हैशटैग और प्रोफ़ाइल | [TikTok Scraper API](https://www.surfsense.com/tiktok) |
|
||||
| **Google Maps** | स्थानीय बिज़नेस रिसर्च के लिए स्थान, रेटिंग और रिव्यू | [Google Maps Scraper API](https://www.surfsense.com/google-maps) |
|
||||
| **Google Search** | सर्च रिसर्च और मॉनिटरिंग के लिए लाइव SERP | [Google Search API](https://www.surfsense.com/google-search) |
|
||||
| **Indeed** | सार्वजनिक नौकरी लिस्टिंग, सैलरी और पूरे विवरण के साथ, सर्च या कंपनी के अनुसार | [Indeed Scraper API](https://www.surfsense.com/indeed) |
|
||||
| **Amazon** | सार्वजनिक प्रोडक्ट डेटा: कीमतें, रेटिंग, ऑफ़र, विक्रेता और बेस्ट-सेलर रैंक | [Amazon Product API](https://www.surfsense.com/amazon) |
|
||||
| **Web Crawl** | ओपन वेब का कोई भी पेज साफ़-सुथरे, स्ट्रक्चर्ड कंटेंट के रूप में | [Web Crawling API](https://www.surfsense.com/web-crawl) |
|
||||
| **External MCP Connectors** | कोई भी MCP सर्वर अपने एजेंट्स से जोड़ें, Notion, Slack, Jira और अन्य के लिए वन-क्लिक OAuth के साथ | [External MCP Connectors](https://www.surfsense.com/external-mcp-connectors) |
|
||||
|
|
@ -217,7 +218,7 @@ SurfSense एकमात्र ओपन सोर्स प्रोडक्
|
|||
|
||||
| फ़ीचर | Google NotebookLM | SurfSense |
|
||||
|---------|-------------------|-----------|
|
||||
| **एजेंट्स के लिए लाइव वेब डेटा** | नहीं | REST API और MCP के ज़रिए Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search और वेब क्रॉल कनेक्टर |
|
||||
| **एजेंट्स के लिए लाइव वेब डेटा** | नहीं | REST API और MCP के ज़रिए Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed और वेब क्रॉल कनेक्टर |
|
||||
| **MCP सर्वर** | नहीं | हर कनेक्टर नेटिव एजेंट टूल के रूप में उपलब्ध, साथ ही वन-क्लिक OAuth ऐप्स के साथ अपने MCP सर्वर लाने की सुविधा |
|
||||
| **प्रति नोटबुक स्रोत** | 50 (Free) से 600 (Ultra, $249.99/माह) | असीमित |
|
||||
| **नोटबुक की संख्या** | 100 (Free) से 500 (सशुल्क टियर) | असीमित |
|
||||
|
|
|
|||
|
|
@ -22,7 +22,7 @@
|
|||
|
||||
# SurfSense: The Open-Source NotebookLM Alternative for Open Web Research
|
||||
|
||||
SurfSense is the **open-source NotebookLM alternative for AI agents**, an open web research platform with live data connectors. Your agents research the live web with structured data from **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, and any page on the open web**, through one **REST API** or **MCP server**. Scheduled and event-triggered agents turn what they find into briefs and alerts, and a built-in knowledge base keeps every finding searchable with citations.
|
||||
SurfSense is the **open-source NotebookLM alternative for AI agents**, an open web research platform with live data connectors. Your agents research the live web with structured data from **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed, and any page on the open web**, through one **REST API** or **MCP server**. Scheduled and event-triggered agents turn what they find into briefs and alerts, and a built-in knowledge base keeps every finding searchable with citations.
|
||||
|
||||
> [!NOTE]
|
||||
> **📢 A note for our NotebookLM-alternative users**
|
||||
|
|
@ -60,6 +60,7 @@ Ask any capable agent "what is Reddit saying about this product since launch?" o
|
|||
| **TikTok** | Videos, comments, hashtags, and profiles without Research API approval | [TikTok Scraper API](https://www.surfsense.com/tiktok) |
|
||||
| **Google Maps** | Places, ratings, and reviews for local business research | [Google Maps Scraper API](https://www.surfsense.com/google-maps) |
|
||||
| **Google Search** | Live SERPs for search research and monitoring | [Google Search API](https://www.surfsense.com/google-search) |
|
||||
| **Indeed** | Public job postings with salaries and full descriptions, by search or company | [Indeed Scraper API](https://www.surfsense.com/indeed) |
|
||||
| **Amazon** | Public product data: prices, ratings, offers, sellers, and best-seller ranks | [Amazon Product API](https://www.surfsense.com/amazon) |
|
||||
| **Web Crawl** | Any page on the open web as clean, structured content | [Web Crawling API](https://www.surfsense.com/web-crawl) |
|
||||
| **External MCP Connectors** | Bring any MCP server to your agents, with one-click OAuth for Notion, Slack, Jira, and more | [External MCP Connectors](https://www.surfsense.com/external-mcp-connectors) |
|
||||
|
|
@ -217,7 +218,7 @@ Still comparing us as a NotebookLM alternative? Here is the honest breakdown.
|
|||
|
||||
| Feature | Google NotebookLM | SurfSense |
|
||||
|---------|-------------------|-----------|
|
||||
| **Live web data for agents** | No | Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, and web crawl connectors via REST API and MCP |
|
||||
| **Live web data for agents** | No | Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed, and web crawl connectors via REST API and MCP |
|
||||
| **MCP server** | No | Every connector exposed as a native agent tool, plus bring-your-own MCP servers with one-click OAuth apps |
|
||||
| **Sources per Notebook** | 50 (Free) to 600 (Ultra, $249.99/mo) | Unlimited |
|
||||
| **Number of Notebooks** | 100 (Free) to 500 (paid tiers) | Unlimited |
|
||||
|
|
|
|||
|
|
@ -22,7 +22,7 @@
|
|||
|
||||
# SurfSense: A Alternativa Open Source ao NotebookLM para Pesquisa na Web Aberta
|
||||
|
||||
O SurfSense é a **alternativa open source ao NotebookLM para agentes de IA**, uma plataforma de pesquisa na web aberta com conectores de dados ao vivo. Seus agentes pesquisam a web ao vivo com dados estruturados do **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search e de qualquer página da web aberta**, por meio de uma única **API REST** ou de um **servidor MCP**. Agentes agendados ou acionados por eventos transformam o que encontram em relatórios e alertas, e uma base de conhecimento integrada mantém cada descoberta pesquisável, com citações.
|
||||
O SurfSense é a **alternativa open source ao NotebookLM para agentes de IA**, uma plataforma de pesquisa na web aberta com conectores de dados ao vivo. Seus agentes pesquisam a web ao vivo com dados estruturados do **Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed e de qualquer página da web aberta**, por meio de uma única **API REST** ou de um **servidor MCP**. Agentes agendados ou acionados por eventos transformam o que encontram em relatórios e alertas, e uma base de conhecimento integrada mantém cada descoberta pesquisável, com citações.
|
||||
|
||||
> [!NOTE]
|
||||
> **📢 Um recado para nossos usuários que buscavam uma alternativa ao NotebookLM**
|
||||
|
|
@ -60,6 +60,7 @@ Pergunte a qualquer agente capaz "o que o Reddit está dizendo sobre este produt
|
|||
| **TikTok** | Vídeos, comentários, hashtags e perfis sem aprovação da Research API | [TikTok Scraper API](https://www.surfsense.com/tiktok) |
|
||||
| **Google Maps** | Estabelecimentos, avaliações e reviews para pesquisa de negócios locais | [Google Maps Scraper API](https://www.surfsense.com/google-maps) |
|
||||
| **Google Search** | SERPs ao vivo para pesquisa e monitoramento de buscas | [Google Search API](https://www.surfsense.com/google-search) |
|
||||
| **Indeed** | Vagas públicas com salários e descrições completas, por busca ou empresa | [Indeed Scraper API](https://www.surfsense.com/indeed) |
|
||||
| **Amazon** | Dados públicos de produtos: preços, avaliações, ofertas, vendedores e rankings de mais vendidos | [Amazon Product API](https://www.surfsense.com/amazon) |
|
||||
| **Web Crawl** (rastreamento web) | Qualquer página da web aberta como conteúdo limpo e estruturado | [Web Crawling API](https://www.surfsense.com/web-crawl) |
|
||||
| **Conectores MCP externos** | Traga qualquer servidor MCP para seus agentes, com OAuth em um clique para Notion, Slack, Jira e outros | [External MCP Connectors](https://www.surfsense.com/external-mcp-connectors) |
|
||||
|
|
@ -217,7 +218,7 @@ Ainda nos comparando como alternativa ao NotebookLM? Aqui está o comparativo ho
|
|||
|
||||
| Recurso | Google NotebookLM | SurfSense |
|
||||
|---------|-------------------|-----------|
|
||||
| **Dados da web ao vivo para agentes** | Não | Conectores de Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search e rastreamento web via API REST e MCP |
|
||||
| **Dados da web ao vivo para agentes** | Não | Conectores de Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, Indeed e rastreamento web via API REST e MCP |
|
||||
| **Servidor MCP** | Não | Cada conector exposto como ferramenta nativa de agente, além de servidores MCP próprios com apps OAuth em um clique |
|
||||
| **Fontes por Notebook** | 50 (gratuito) a 600 (Ultra, US$ 249,99/mês) | Ilimitadas |
|
||||
| **Número de Notebooks** | 100 (gratuito) a 500 (planos pagos) | Ilimitado |
|
||||
|
|
|
|||
|
|
@ -22,7 +22,7 @@
|
|||
|
||||
# SurfSense:面向开放网络研究的开源 NotebookLM 替代品
|
||||
|
||||
SurfSense 是**面向 AI 智能体的开源 NotebookLM 替代品**,一个配备实时数据连接器的开放网络研究平台。你的智能体可以通过一个 **REST API** 或 **MCP 服务器**,利用来自 **Reddit、YouTube、Instagram、TikTok、Amazon、Google Maps、Google Search 以及开放网络上任意页面**的结构化数据研究实时网络。定时和事件触发的智能体会把发现的内容转化为简报和预警,内置的知识库则让每一条发现都可搜索、可引用。
|
||||
SurfSense 是**面向 AI 智能体的开源 NotebookLM 替代品**,一个配备实时数据连接器的开放网络研究平台。你的智能体可以通过一个 **REST API** 或 **MCP 服务器**,利用来自 **Reddit、YouTube、Instagram、TikTok、Amazon、Google Maps、Google Search、Indeed 以及开放网络上任意页面**的结构化数据研究实时网络。定时和事件触发的智能体会把发现的内容转化为简报和预警,内置的知识库则让每一条发现都可搜索、可引用。
|
||||
|
||||
> [!NOTE]
|
||||
> **📢 致我们的 NotebookLM 替代品用户**
|
||||
|
|
@ -60,6 +60,7 @@ SurfSense 是**面向 AI 智能体的开源 NotebookLM 替代品**,一个配
|
|||
| **TikTok** | 视频、评论、话题标签和主页,无需 Research API 审批 | [TikTok Scraper API](https://www.surfsense.com/tiktok) |
|
||||
| **Google Maps** | 地点、评分和评论,用于本地商户研究 | [Google Maps Scraper API](https://www.surfsense.com/google-maps) |
|
||||
| **Google Search** | 实时搜索结果页,用于搜索研究和监控 | [Google Search API](https://www.surfsense.com/google-search) |
|
||||
| **Indeed** | 公开职位信息,含薪资与完整职位描述,按搜索或公司抓取 | [Indeed Scraper API](https://www.surfsense.com/indeed) |
|
||||
| **Amazon** | 公开商品数据:价格、评分、报价、卖家和畅销榜排名 | [Amazon Product API](https://www.surfsense.com/amazon) |
|
||||
| **Web Crawl** | 把开放网络上的任意页面转为干净、结构化的内容 | [Web Crawling API](https://www.surfsense.com/web-crawl) |
|
||||
| **外部 MCP 连接器** | 将任意 MCP 服务器接入你的智能体,Notion、Slack、Jira 等支持一键 OAuth | [External MCP Connectors](https://www.surfsense.com/external-mcp-connectors) |
|
||||
|
|
@ -217,7 +218,7 @@ SurfSense 是唯一一款把面向人的 NotebookLM 式研究工作区与面向
|
|||
|
||||
| 功能 | Google NotebookLM | SurfSense |
|
||||
|---------|-------------------|-----------|
|
||||
| **面向智能体的实时网络数据** | 无 | 通过 REST API 和 MCP 提供 Reddit、YouTube、Instagram、TikTok、Amazon、Google Maps、Google Search 和网页爬取连接器 |
|
||||
| **面向智能体的实时网络数据** | 无 | 通过 REST API 和 MCP 提供 Reddit、YouTube、Instagram、TikTok、Amazon、Google Maps、Google Search、Indeed 和网页爬取连接器 |
|
||||
| **MCP 服务器** | 无 | 每个连接器都作为原生智能体工具暴露,还可自带 MCP 服务器并使用一键 OAuth 应用 |
|
||||
| **每个笔记本的来源数** | 50 个(免费版)至 600 个(Ultra 版,249.99 美元/月) | 无限制 |
|
||||
| **笔记本数量** | 100 个(免费版)至 500 个(付费档位) | 无限制 |
|
||||
|
|
|
|||
|
|
@ -458,6 +458,7 @@ SURFSENSE_ENABLE_DOOM_LOOP=true
|
|||
# TIKTOK_MICROS_PER_VIDEO=3500
|
||||
# TIKTOK_MICROS_PER_USER=2500
|
||||
# TIKTOK_MICROS_PER_COMMENT=1500
|
||||
# INDEED_SCRAPE_MICROS_PER_JOB=3500
|
||||
|
||||
# Safety ceiling on per-call premium reservation, in micro-USD ($1.00 default).
|
||||
# QUOTA_MAX_RESERVE_MICROS=1000000
|
||||
|
|
|
|||
|
|
@ -298,6 +298,7 @@ MICROS_PER_PAGE=1000
|
|||
# TIKTOK_MICROS_PER_VIDEO=3500
|
||||
# TIKTOK_MICROS_PER_USER=2500
|
||||
# TIKTOK_MICROS_PER_COMMENT=1500
|
||||
# INDEED_SCRAPE_MICROS_PER_JOB=3500
|
||||
# Browser-listing retries when a feed is empty (profile feed is withheld from
|
||||
# flagged IPs; each retry draws a fresh rotating exit IP). Set to 1 for a static IP.
|
||||
# TIKTOK_LISTING_MAX_ATTEMPTS=3
|
||||
|
|
|
|||
|
|
@ -36,6 +36,7 @@ SUBAGENT_TO_REQUIRED_CONNECTOR_MAP: dict[str, frozenset[str]] = {
|
|||
"youtube": frozenset(),
|
||||
"google_maps": frozenset(),
|
||||
"google_search": frozenset(),
|
||||
"indeed": frozenset(),
|
||||
"reddit": frozenset(),
|
||||
"instagram": frozenset(),
|
||||
"tiktok": frozenset(),
|
||||
|
|
|
|||
|
|
@ -0,0 +1 @@
|
|||
"""``indeed`` builtin subagent: structured Indeed job postings."""
|
||||
|
|
@ -0,0 +1,43 @@
|
|||
"""``indeed`` route: ``SurfSenseSubagentSpec`` builder for deepagents."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
from langchain_core.language_models import BaseChatModel
|
||||
from langchain_core.tools import BaseTool
|
||||
|
||||
from app.agents.chat.multi_agent_chat.subagents.shared.md_file_reader import (
|
||||
read_md_file,
|
||||
)
|
||||
from app.agents.chat.multi_agent_chat.subagents.shared.spec import SurfSenseSubagentSpec
|
||||
from app.agents.chat.multi_agent_chat.subagents.shared.subagent_builder import (
|
||||
pack_subagent,
|
||||
)
|
||||
|
||||
from .tools.index import NAME, RULESET, load_tools
|
||||
|
||||
|
||||
def build_subagent(
|
||||
*,
|
||||
dependencies: dict[str, Any],
|
||||
model: BaseChatModel | None = None,
|
||||
middleware_stack: dict[str, Any] | None = None,
|
||||
mcp_tools: list[BaseTool] | None = None,
|
||||
) -> SurfSenseSubagentSpec:
|
||||
tools = [*load_tools(dependencies=dependencies), *(mcp_tools or [])]
|
||||
description = (
|
||||
read_md_file(__package__, "description").strip()
|
||||
or "Pulls structured job postings from Indeed search and company pages."
|
||||
)
|
||||
system_prompt = read_md_file(__package__, "system_prompt").strip()
|
||||
return pack_subagent(
|
||||
name=NAME,
|
||||
description=description,
|
||||
system_prompt=system_prompt,
|
||||
tools=tools,
|
||||
ruleset=RULESET,
|
||||
dependencies=dependencies,
|
||||
model=model,
|
||||
middleware_stack=middleware_stack,
|
||||
)
|
||||
|
|
@ -0,0 +1,2 @@
|
|||
Indeed jobs specialist: pulls structured job postings — title, company, location, salary (range, currency, period), job types, benefits, remote/hybrid flag, posting age, apply URL, and the full job description. Discovers jobs by search query (with country, location, radius, job type, experience level, remote/hybrid, and date-posted filters), or scrapes a known Indeed search, company jobs, or single job URL as-is. Optionally fetches each job's detail page for the full description.
|
||||
Use whenever the task is to find or gather job listings from Indeed — openings for a role, hiring at a company, salaries for a title in a location, or remote roles in a field. Triggers include "find jobs for X", "who is hiring X", "data analyst jobs in Y", "remote X roles", and scraping a specific Indeed URL. Not for general web pages (use the web crawling specialist), Google results (use the Google Search specialist), or other job boards.
|
||||
|
|
@ -0,0 +1,65 @@
|
|||
You are the SurfSense Indeed sub-agent.
|
||||
You receive delegated instructions from a supervisor agent and return structured results for supervisor synthesis.
|
||||
|
||||
<goal>
|
||||
Answer the delegated question from live Indeed job data gathered with your verb, comparing against earlier results already in this conversation when the task calls for it.
|
||||
</goal>
|
||||
|
||||
<available_tools>
|
||||
- `indeed_scrape`
|
||||
- `read_run` / `search_run` (free readers for stored scrape output)
|
||||
</available_tools>
|
||||
|
||||
<playbook>
|
||||
- Finding jobs for a role: call `indeed_scrape` with `search_queries`; narrow with `location`, `country`, `job_type`, `level`, `remote`, `radius`, and `from_days`.
|
||||
- Scraping a specific Indeed URL: pass a search, company jobs, or single job URL in `urls`.
|
||||
- Full descriptions: set `scrape_job_details=true` to fetch each job's detail page (slower: one extra load per job). Leave it false when the listing snippet is enough.
|
||||
- Cost model: Indeed is latency-bound. A cold session spends minutes solving Cloudflare, and each query returns only its first page (~15 jobs) — anonymous pagination is gated, so `max_items` above ~15/query buys nothing. Every extra phrasing adds wait, not depth.
|
||||
- Default to ONE focused query. Only add phrasings when the role genuinely needs variety, and then put at most 2–3 in a SINGLE call's `search_queries` (they reuse one warmed session); never make separate `indeed_scrape` calls, which each pay the cold-start cost.
|
||||
- Prefer returning the first page's on-topic hits promptly over exhaustive coverage. If a single call under-delivers against a large requested N, return `status=partial` noting the ~one-page-per-query ceiling — do not chase it with more calls.
|
||||
<include snippet="run_reader"/>
|
||||
- Comparison requests: pull the current results, compare against prior values already in this conversation's earlier tool results, and report concrete deltas (added, removed, salary/rank changes).
|
||||
</playbook>
|
||||
|
||||
<tool_policy>
|
||||
- Use only tools in `<available_tools>`.
|
||||
- Report only results present in the tool output. Never invent titles, companies, salaries, locations, or description text.
|
||||
</tool_policy>
|
||||
|
||||
<out_of_scope>
|
||||
- Do not read arbitrary web pages — that belongs to the web crawling specialist.
|
||||
- Do not generate deliverables or perform connector mutations; return findings for the supervisor to act on.
|
||||
- Google results belong to the Google Search specialist; other job boards are out of scope.
|
||||
</out_of_scope>
|
||||
|
||||
<safety>
|
||||
- Report uncertainty explicitly when evidence is incomplete or conflicting.
|
||||
- Never present unverified claims as facts.
|
||||
</safety>
|
||||
|
||||
<failure_policy>
|
||||
- Underspecified request — no usable query or URL — return `status=blocked` with the missing fields.
|
||||
- Tool failure: return `status=error` with a concise recovery `next_step`.
|
||||
- No useful evidence: return `status=blocked` with a narrower query or the scope you still need.
|
||||
</failure_policy>
|
||||
|
||||
<output_contract>
|
||||
Return **only** one JSON object (no markdown/prose):
|
||||
{
|
||||
"status": "success" | "partial" | "blocked" | "error",
|
||||
"action_summary": string,
|
||||
"evidence": {
|
||||
"findings": string[],
|
||||
"sources": string[],
|
||||
"confidence": "high" | "medium" | "low"
|
||||
},
|
||||
"next_step": string | null,
|
||||
"missing_fields": string[] | null,
|
||||
"assumptions": string[] | null
|
||||
}
|
||||
<include snippet="output_contract_base"/>
|
||||
Route-specific rules:
|
||||
- `evidence.findings`: one entry per distinct job or delta — a single sentence each; do not paste raw payloads. Max 10 entries, unless the delegated task asks for N items: then return up to N (each backed by a real scraped result, never padded).
|
||||
- `evidence.sources`: one Indeed job URL per finding when applicable, same cap as findings. List each URL once.
|
||||
</output_contract>
|
||||
</output>
|
||||
|
|
@ -0,0 +1,27 @@
|
|||
"""``indeed`` sub-agent tools: the Indeed scrape capability verb."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any
|
||||
|
||||
from langchain_core.tools import BaseTool
|
||||
|
||||
from app.agents.chat.multi_agent_chat.shared.permissions import Ruleset
|
||||
from app.capabilities.core.access.agent import build_capability_tools
|
||||
from app.capabilities.indeed.scrape.definition import INDEED_SCRAPE
|
||||
|
||||
NAME = "indeed"
|
||||
|
||||
RULESET = Ruleset(origin=NAME, rules=[])
|
||||
|
||||
_CI_VERBS = [INDEED_SCRAPE]
|
||||
|
||||
|
||||
def load_tools(
|
||||
*, dependencies: dict[str, Any] | None = None, **kwargs: Any
|
||||
) -> list[BaseTool]:
|
||||
d = {**(dependencies or {}), **kwargs}
|
||||
return build_capability_tools(
|
||||
workspace_id=d.get("workspace_id"),
|
||||
capabilities=_CI_VERBS,
|
||||
)
|
||||
|
|
@ -24,6 +24,9 @@ from app.agents.chat.multi_agent_chat.subagents.builtins.google_maps.agent impor
|
|||
from app.agents.chat.multi_agent_chat.subagents.builtins.google_search.agent import (
|
||||
build_subagent as build_google_search_subagent,
|
||||
)
|
||||
from app.agents.chat.multi_agent_chat.subagents.builtins.indeed.agent import (
|
||||
build_subagent as build_indeed_subagent,
|
||||
)
|
||||
from app.agents.chat.multi_agent_chat.subagents.builtins.instagram.agent import (
|
||||
build_subagent as build_instagram_subagent,
|
||||
)
|
||||
|
|
@ -89,6 +92,7 @@ SUBAGENT_BUILDERS_BY_NAME: dict[str, SubagentBuilder] = {
|
|||
"google_drive": build_google_drive_subagent,
|
||||
"google_maps": build_google_maps_subagent,
|
||||
"google_search": build_google_search_subagent,
|
||||
"indeed": build_indeed_subagent,
|
||||
"instagram": build_instagram_subagent,
|
||||
"knowledge_base": build_knowledge_base_subagent,
|
||||
"mcp_discovery": build_mcp_discovery_subagent,
|
||||
|
|
|
|||
|
|
@ -41,6 +41,7 @@ _PLATFORM_RATE_KEYS: dict[BillingUnit, str] = {
|
|||
BillingUnit.TIKTOK_VIDEO: "TIKTOK_MICROS_PER_VIDEO",
|
||||
BillingUnit.TIKTOK_USER: "TIKTOK_MICROS_PER_USER",
|
||||
BillingUnit.TIKTOK_COMMENT: "TIKTOK_MICROS_PER_COMMENT",
|
||||
BillingUnit.INDEED_JOB: "INDEED_SCRAPE_MICROS_PER_JOB",
|
||||
}
|
||||
|
||||
|
||||
|
|
@ -63,6 +64,7 @@ _UNIT_NOUNS: dict[BillingUnit, str] = {
|
|||
BillingUnit.TIKTOK_VIDEO: "video",
|
||||
BillingUnit.TIKTOK_USER: "profile",
|
||||
BillingUnit.TIKTOK_COMMENT: "comment",
|
||||
BillingUnit.INDEED_JOB: "job",
|
||||
}
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -31,6 +31,7 @@ class BillingUnit(StrEnum):
|
|||
TIKTOK_VIDEO = "tiktok_video"
|
||||
TIKTOK_USER = "tiktok_user"
|
||||
TIKTOK_COMMENT = "tiktok_comment"
|
||||
INDEED_JOB = "indeed_job"
|
||||
|
||||
|
||||
class BillableInput(Protocol):
|
||||
|
|
|
|||
5
surfsense_backend/app/capabilities/indeed/__init__.py
Normal file
5
surfsense_backend/app/capabilities/indeed/__init__.py
Normal file
|
|
@ -0,0 +1,5 @@
|
|||
"""``indeed.*`` namespace: platform-native Indeed data verbs."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from app.capabilities.indeed.scrape import definition as _scrape # noqa: F401
|
||||
|
|
@ -0,0 +1,3 @@
|
|||
"""``indeed.scrape`` verb: Indeed search / company URLs → job postings."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
|
@ -0,0 +1,23 @@
|
|||
"""``indeed.scrape`` capability registration (billed per job; see config
|
||||
``INDEED_SCRAPE_MICROS_PER_JOB``)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from app.capabilities.core import BillingUnit, Capability, register_capability
|
||||
from app.capabilities.indeed.scrape.executor import build_scrape_executor
|
||||
from app.capabilities.indeed.scrape.schemas import ScrapeInput, ScrapeOutput
|
||||
|
||||
INDEED_SCRAPE = Capability(
|
||||
name="indeed.scrape",
|
||||
description=(
|
||||
"Scrape public Indeed job postings, including title, company, location, "
|
||||
"salary, and description. Use urls or search_queries."
|
||||
),
|
||||
input_schema=ScrapeInput,
|
||||
output_schema=ScrapeOutput,
|
||||
executor=build_scrape_executor(),
|
||||
billing_unit=BillingUnit.INDEED_JOB,
|
||||
docs_url="/docs/connectors/native/indeed",
|
||||
)
|
||||
|
||||
register_capability(INDEED_SCRAPE)
|
||||
56
surfsense_backend/app/capabilities/indeed/scrape/executor.py
Normal file
56
surfsense_backend/app/capabilities/indeed/scrape/executor.py
Normal file
|
|
@ -0,0 +1,56 @@
|
|||
"""``indeed.scrape`` executor: verb input → scraper → Indeed job items."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from collections.abc import Awaitable, Callable
|
||||
|
||||
from app.capabilities.core import Executor
|
||||
from app.capabilities.core.progress import emit_progress
|
||||
from app.capabilities.indeed.scrape.schemas import ScrapeInput, ScrapeOutput
|
||||
from app.exceptions import ForbiddenError
|
||||
from app.proprietary.platforms.indeed_jobs import (
|
||||
IndeedAccessBlockedError,
|
||||
IndeedScrapeInput,
|
||||
scrape_indeed,
|
||||
)
|
||||
|
||||
ScrapeFn = Callable[..., Awaitable[list[dict]]]
|
||||
|
||||
|
||||
def build_scrape_executor(scrape_fn: ScrapeFn | None = None) -> Executor:
|
||||
"""Bind the executor to a scraper fn (defaults to the proprietary actor)."""
|
||||
scrape_fn = scrape_fn or scrape_indeed
|
||||
|
||||
async def execute(payload: ScrapeInput) -> ScrapeOutput:
|
||||
actor_input = IndeedScrapeInput(
|
||||
startUrls=[{"url": url} for url in payload.urls],
|
||||
queries=payload.search_queries,
|
||||
country=payload.country,
|
||||
location=payload.location,
|
||||
radius=payload.radius,
|
||||
jobType=payload.job_type,
|
||||
level=payload.level,
|
||||
remote=payload.remote,
|
||||
fromDays=payload.from_days,
|
||||
sort=payload.sort,
|
||||
scrapeJobDetails=payload.scrape_job_details,
|
||||
maxItems=payload.max_items,
|
||||
maxItemsPerQuery=payload.max_items_per_query,
|
||||
)
|
||||
emit_progress(
|
||||
"starting", "Resolving Indeed targets", total=payload.max_items, unit="job"
|
||||
)
|
||||
try:
|
||||
items = await scrape_fn(actor_input, limit=payload.max_items)
|
||||
except IndeedAccessBlockedError as exc:
|
||||
# Anonymous-only scraper; a hard block can't be retried with creds.
|
||||
raise ForbiddenError(
|
||||
f"Indeed refused anonymous access: {exc}",
|
||||
code="INDEED_ACCESS_BLOCKED",
|
||||
) from exc
|
||||
emit_progress(
|
||||
"done", f"Scraped {len(items)} job(s)", current=len(items), unit="job"
|
||||
)
|
||||
return ScrapeOutput(items=items)
|
||||
|
||||
return execute
|
||||
119
surfsense_backend/app/capabilities/indeed/scrape/schemas.py
Normal file
119
surfsense_backend/app/capabilities/indeed/scrape/schemas.py
Normal file
|
|
@ -0,0 +1,119 @@
|
|||
"""``indeed.scrape`` I/O contracts.
|
||||
|
||||
A lean, agent-friendly surface over ``IndeedScrapeInput``
|
||||
(``app/proprietary/platforms/indeed_jobs``). The executor maps this to the full
|
||||
scraper input; the scraper's ``IndeedItem`` is reused verbatim as the output
|
||||
element.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from pydantic import BaseModel, Field, model_validator
|
||||
|
||||
from app.proprietary.platforms.indeed_jobs import IndeedItem
|
||||
from app.proprietary.platforms.indeed_jobs.schemas import (
|
||||
IndeedJobType,
|
||||
IndeedLevel,
|
||||
IndeedRemote,
|
||||
IndeedSort,
|
||||
)
|
||||
|
||||
MAX_INDEED_SOURCES = 20
|
||||
"""Per-call cap on urls + search_queries: bounds a synchronous request's fan-out."""
|
||||
|
||||
MAX_INDEED_ITEMS = 100
|
||||
"""Hard ceiling on jobs returned per call, regardless of the per-query caps."""
|
||||
|
||||
|
||||
class ScrapeInput(BaseModel):
|
||||
urls: list[str] = Field(
|
||||
default_factory=list,
|
||||
max_length=MAX_INDEED_SOURCES,
|
||||
description=(
|
||||
"Indeed URLs to scrape: a search page (/jobs?q=&l=) or a company "
|
||||
"jobs page (/cmp/<slug>/jobs). Provide these OR search_queries "
|
||||
"(at least one source is required)."
|
||||
),
|
||||
)
|
||||
search_queries: list[str] = Field(
|
||||
default_factory=list,
|
||||
max_length=MAX_INDEED_SOURCES,
|
||||
description=(
|
||||
"Job search terms; each returns up to max_items_per_query results, "
|
||||
"shaped by country/location/job_type/etc."
|
||||
),
|
||||
)
|
||||
country: str = Field(
|
||||
default="us",
|
||||
description="Country code selecting the Indeed domain, e.g. 'us', 'gb', 'de'.",
|
||||
)
|
||||
location: str | None = Field(
|
||||
default=None,
|
||||
description="Where to search, e.g. 'Remote', 'New York, NY'.",
|
||||
)
|
||||
radius: int | None = Field(
|
||||
default=None,
|
||||
description="Search radius in miles/km around location.",
|
||||
)
|
||||
job_type: IndeedJobType | None = Field(
|
||||
default=None,
|
||||
description="Employment type filter: fulltime, parttime, contract, etc.",
|
||||
)
|
||||
level: IndeedLevel | None = Field(
|
||||
default=None,
|
||||
description="Experience level filter: entry_level, mid_level, senior_level.",
|
||||
)
|
||||
remote: IndeedRemote | None = Field(
|
||||
default=None,
|
||||
description="Work model filter: remote or hybrid.",
|
||||
)
|
||||
from_days: int | None = Field(
|
||||
default=None,
|
||||
description="Only return jobs posted within the last N days.",
|
||||
)
|
||||
sort: IndeedSort = Field(
|
||||
default="relevance",
|
||||
description="Result ordering: relevance or date.",
|
||||
)
|
||||
scrape_job_details: bool = Field(
|
||||
default=False,
|
||||
description=(
|
||||
"Fetch each job's detail page for the full description (slower: one "
|
||||
"extra page load per job)."
|
||||
),
|
||||
)
|
||||
max_items: int = Field(
|
||||
default=25,
|
||||
ge=1,
|
||||
le=MAX_INDEED_ITEMS,
|
||||
description="Max total jobs to return across all sources.",
|
||||
)
|
||||
max_items_per_query: int = Field(
|
||||
default=25,
|
||||
ge=0,
|
||||
description="Max jobs to pull per search/company target.",
|
||||
)
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _require_a_source(self) -> ScrapeInput:
|
||||
if not self.urls and not self.search_queries:
|
||||
raise ValueError("Provide at least one of 'urls' or 'search_queries'.")
|
||||
return self
|
||||
|
||||
@property
|
||||
def estimated_units(self) -> int:
|
||||
"""Worst-case billable jobs for the pre-flight gate: ``max_items`` is a
|
||||
hard cross-source ceiling (le=100), so no call can exceed it."""
|
||||
return self.max_items
|
||||
|
||||
|
||||
class ScrapeOutput(BaseModel):
|
||||
items: list[IndeedItem] = Field(
|
||||
default_factory=list,
|
||||
description="One item per job posting, in emission order.",
|
||||
)
|
||||
|
||||
@property
|
||||
def billable_units(self) -> int:
|
||||
"""One returned job = one billable unit."""
|
||||
return len(self.items)
|
||||
|
|
@ -736,6 +736,11 @@ class Config:
|
|||
# Comments are the cheapest per-item TikTok data, matching the per-comment
|
||||
# market (and YouTube's comment meter).
|
||||
TIKTOK_MICROS_PER_COMMENT = int(os.getenv("TIKTOK_MICROS_PER_COMMENT", "1500"))
|
||||
# Warmed-browser listings put Indeed on par with the other browser-driven
|
||||
# scrapers (Reddit, Instagram) rather than the cheaper API-backed meters.
|
||||
INDEED_SCRAPE_MICROS_PER_JOB = int(
|
||||
os.getenv("INDEED_SCRAPE_MICROS_PER_JOB", "3500")
|
||||
)
|
||||
# Retry an empty listing draw on a fresh rotating IP. Set to 1 for a static
|
||||
# proxy, where every retry re-hits the same exit.
|
||||
TIKTOK_LISTING_MAX_ATTEMPTS = int(os.getenv("TIKTOK_LISTING_MAX_ATTEMPTS", "3"))
|
||||
|
|
|
|||
|
|
@ -0,0 +1,13 @@
|
|||
"""Platform-native Indeed jobs scraper (anonymous, warmed browser session)."""
|
||||
|
||||
from .fetch import IndeedAccessBlockedError
|
||||
from .schemas import IndeedItem, IndeedScrapeInput
|
||||
from .scraper import iter_indeed, scrape_indeed
|
||||
|
||||
__all__ = [
|
||||
"IndeedAccessBlockedError",
|
||||
"IndeedItem",
|
||||
"IndeedScrapeInput",
|
||||
"iter_indeed",
|
||||
"scrape_indeed",
|
||||
]
|
||||
188
surfsense_backend/app/proprietary/platforms/indeed_jobs/fetch.py
Normal file
188
surfsense_backend/app/proprietary/platforms/indeed_jobs/fetch.py
Normal file
|
|
@ -0,0 +1,188 @@
|
|||
"""Browser-session fetch seam for the Indeed scraper.
|
||||
|
||||
Indeed fronts its origin with Cloudflare plus an anonymous-bot check that bounces
|
||||
cold sessions to ``secure.indeed.com/auth``. The working recipe: a persistent
|
||||
camoufox session that solves Cloudflare, warms on the domain home page, then
|
||||
navigates to ``/jobs`` in the same context so the clearance carries.
|
||||
|
||||
:class:`IndeedSession` warms per domain once, retries a blocked page on a fresh
|
||||
residential IP, and caps each navigation with a hard timeout so a stuck solve
|
||||
can't stall a run. All egress is through the residential proxy.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import logging
|
||||
from collections.abc import Awaitable, Callable
|
||||
from contextlib import asynccontextmanager, suppress
|
||||
from datetime import UTC, datetime
|
||||
from typing import Any, Protocol
|
||||
from urllib.parse import urlparse
|
||||
|
||||
from app.utils.proxy import get_proxy_url
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
|
||||
class IndeedAccessBlockedError(RuntimeError):
|
||||
"""Every rotated IP was bounced to Indeed's security wall."""
|
||||
|
||||
|
||||
# Per navigation; a stuck Cloudflare solve otherwise hangs the whole run.
|
||||
_PAGE_TIMEOUT_S = 75.0
|
||||
# Browser-internal timeout; kept above the page timeout so ours fires first.
|
||||
_SESSION_TIMEOUT_MS = 90_000
|
||||
_MAX_ROTATIONS = 3
|
||||
|
||||
# Markers of a Cloudflare / security-check interstitial served instead of jobs.
|
||||
_BLOCK_MARKERS = (
|
||||
"secure.indeed.com",
|
||||
"bot-detection",
|
||||
"security check",
|
||||
"challenge-platform",
|
||||
"just a moment",
|
||||
"verify you are human",
|
||||
"hcaptcha",
|
||||
)
|
||||
|
||||
|
||||
def now_iso() -> str:
|
||||
"""UTC timestamp in the millisecond ISO shape used by scraper output."""
|
||||
return datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%S.%f")[:-3] + "Z"
|
||||
|
||||
|
||||
class _Session(Protocol):
|
||||
"""Minimal browser-session surface used here (real or fake)."""
|
||||
|
||||
async def start(self) -> Any: ...
|
||||
async def fetch(self, url: str, **kwargs: Any) -> Any: ...
|
||||
async def close(self) -> Any: ...
|
||||
|
||||
|
||||
def _default_session_factory() -> _Session:
|
||||
"""Build a proxied, Cloudflare-solving camoufox session.
|
||||
|
||||
``disable_resources`` skips images/fonts/media; job data is inline in the
|
||||
document, so this only trims bandwidth.
|
||||
"""
|
||||
from scrapling.fetchers import AsyncStealthySession
|
||||
|
||||
return AsyncStealthySession(
|
||||
headless=True,
|
||||
solve_cloudflare=True,
|
||||
network_idle=True,
|
||||
block_webrtc=True,
|
||||
disable_resources=True,
|
||||
timeout=_SESSION_TIMEOUT_MS,
|
||||
proxy=get_proxy_url(),
|
||||
)
|
||||
|
||||
|
||||
def _html(page: Any) -> str:
|
||||
"""Best-effort HTML body across scrapling response shapes."""
|
||||
for attr in ("html_content", "body", "text"):
|
||||
val = getattr(page, attr, None)
|
||||
if isinstance(val, bytes):
|
||||
val = val.decode("utf-8", "replace")
|
||||
if isinstance(val, str) and val:
|
||||
return val
|
||||
return ""
|
||||
|
||||
|
||||
def _looks_blocked(html: str, final_url: str) -> bool:
|
||||
"""Whether a response is an interstitial rather than a real page."""
|
||||
if not html:
|
||||
return True
|
||||
haystack = (final_url + " " + html[:6000]).lower()
|
||||
return any(marker in haystack for marker in _BLOCK_MARKERS)
|
||||
|
||||
|
||||
class IndeedSession:
|
||||
"""One warmed browser session that rotates its exit IP when blocked."""
|
||||
|
||||
def __init__(
|
||||
self, session_factory: Callable[[], _Session] = _default_session_factory
|
||||
) -> None:
|
||||
self._factory = session_factory
|
||||
self._session: _Session | None = None
|
||||
self._warmed: set[str] = set()
|
||||
self.rotations = 0
|
||||
|
||||
async def start(self) -> None:
|
||||
self._session = self._factory()
|
||||
await self._session.start()
|
||||
|
||||
async def close(self) -> None:
|
||||
if self._session is not None:
|
||||
with suppress(Exception):
|
||||
await self._session.close()
|
||||
self._session = None
|
||||
self._warmed.clear()
|
||||
|
||||
async def _rotate(self) -> None:
|
||||
"""Drop the session for a fresh exit IP; clears warmed domains."""
|
||||
await self.close()
|
||||
self.rotations += 1
|
||||
await self.start()
|
||||
logger.info("[indeed] rotated session (rotation #%d)", self.rotations)
|
||||
|
||||
async def _timed_fetch(self, url: str, **kwargs: Any) -> Any:
|
||||
assert self._session is not None
|
||||
coro: Awaitable[Any] = self._session.fetch(url, **kwargs)
|
||||
return await asyncio.wait_for(coro, timeout=_PAGE_TIMEOUT_S)
|
||||
|
||||
async def _ensure_warm(self, domain: str) -> None:
|
||||
"""Land on the domain home with a Google referer before scraping it."""
|
||||
if domain in self._warmed:
|
||||
return
|
||||
with suppress(Exception):
|
||||
await self._timed_fetch(f"https://{domain}/", google_search=True)
|
||||
self._warmed.add(domain)
|
||||
|
||||
async def fetch_html(self, url: str, *, max_rotations: int | None = None) -> str:
|
||||
"""Return a search/company/job page's HTML through the warmed session.
|
||||
|
||||
Rotates the IP and re-warms on a security-wall bounce or timeout; raises
|
||||
:class:`IndeedAccessBlockedError` once rotations are exhausted. ``max_rotations``
|
||||
overrides the default budget: pass ``0`` to fail fast on a systematically
|
||||
gated page (e.g. anonymous pagination) instead of burning rotations on a
|
||||
block no fresh IP will clear.
|
||||
"""
|
||||
if self._session is None:
|
||||
await self.start()
|
||||
budget = _MAX_ROTATIONS if max_rotations is None else max_rotations
|
||||
domain = urlparse(url).hostname or "www.indeed.com"
|
||||
attempt = 0
|
||||
while True:
|
||||
try:
|
||||
await self._ensure_warm(domain)
|
||||
page = await self._timed_fetch(url)
|
||||
html = _html(page)
|
||||
if not _looks_blocked(html, str(getattr(page, "url", "") or "")):
|
||||
return html
|
||||
logger.info("[indeed] blocked on %s", url)
|
||||
except TimeoutError:
|
||||
logger.warning("[indeed] fetch timed out on %s", url)
|
||||
except Exception as e:
|
||||
logger.warning("[indeed] fetch failed on %s: %s", url, e)
|
||||
|
||||
if attempt >= budget:
|
||||
raise IndeedAccessBlockedError(
|
||||
f"Indeed refused {url} after {attempt + 1} attempt(s)"
|
||||
)
|
||||
attempt += 1
|
||||
await self._rotate()
|
||||
|
||||
|
||||
@asynccontextmanager
|
||||
async def open_session(
|
||||
session_factory: Callable[[], _Session] = _default_session_factory,
|
||||
):
|
||||
"""Open an :class:`IndeedSession` and guarantee teardown."""
|
||||
session = IndeedSession(session_factory)
|
||||
await session.start()
|
||||
try:
|
||||
yield session
|
||||
finally:
|
||||
await session.close()
|
||||
|
|
@ -0,0 +1,327 @@
|
|||
"""Pure HTML/JSON -> item mapping for the Indeed scraper.
|
||||
|
||||
I/O-free and deterministic so it can be unit-tested against captured fixtures;
|
||||
the orchestrator stamps ``scrapedAt``.
|
||||
|
||||
Indeed embeds job data as a JS assignment::
|
||||
|
||||
window.mosaic.providerData["mosaic-provider-jobcards"]={"metaData":{...}};
|
||||
|
||||
The literal recurs hundreds of times in the page, so :func:`extract_jobcards_blob`
|
||||
anchors on the assignment (``...]=``) and brace-matches the balanced object.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from datetime import UTC, datetime
|
||||
from html import unescape
|
||||
from re import sub as _re_sub
|
||||
from typing import Any
|
||||
|
||||
_DEFAULT_BASE = "https://www.indeed.com"
|
||||
|
||||
_JOBCARDS_ANCHOR = 'window.mosaic.providerData["mosaic-provider-jobcards"]='
|
||||
|
||||
# A /viewjob page carries the posting model in ``window._rootProps`` (JSON,
|
||||
# under ``preloadedVJData``); older pages inlined it as ``window._initialData``.
|
||||
_ROOT_PROPS_ANCHOR = "window._rootProps"
|
||||
_ROOT_PROPS_KEY = "preloadedVJData"
|
||||
_INITIAL_DATA_ANCHOR = "window._initialData"
|
||||
|
||||
# Indeed's extractedSalary.type -> our SalaryPeriod.
|
||||
_SALARY_PERIODS = {
|
||||
"HOURLY": "hour",
|
||||
"DAILY": "day",
|
||||
"WEEKLY": "week",
|
||||
"MONTHLY": "month",
|
||||
"YEARLY": "year",
|
||||
}
|
||||
|
||||
|
||||
def _brace_match(text: str, start: int) -> str | None:
|
||||
"""Return the balanced ``{...}``/``[...]`` blob at ``text[start]``, quote-aware."""
|
||||
open_ch = text[start] if start < len(text) else ""
|
||||
close_ch = {"[": "]", "{": "}"}.get(open_ch)
|
||||
if close_ch is None:
|
||||
return None
|
||||
depth = 0
|
||||
i = start
|
||||
n = len(text)
|
||||
while i < n:
|
||||
ch = text[i]
|
||||
if ch == open_ch:
|
||||
depth += 1
|
||||
elif ch == close_ch:
|
||||
depth -= 1
|
||||
if depth == 0:
|
||||
return text[start : i + 1]
|
||||
elif ch == '"':
|
||||
i += 1
|
||||
while i < n and text[i] != '"':
|
||||
if text[i] == "\\":
|
||||
i += 1
|
||||
i += 1
|
||||
i += 1
|
||||
return None
|
||||
|
||||
|
||||
def extract_jobcards_blob(html: str) -> dict | None:
|
||||
"""Decode the ``mosaic-provider-jobcards`` assignment, or ``None`` if absent."""
|
||||
import json
|
||||
|
||||
idx = html.find(_JOBCARDS_ANCHOR)
|
||||
if idx == -1:
|
||||
return None
|
||||
brace = html.find("{", idx + len(_JOBCARDS_ANCHOR))
|
||||
if brace == -1:
|
||||
return None
|
||||
blob = _brace_match(html, brace)
|
||||
if not blob:
|
||||
return None
|
||||
try:
|
||||
data = json.loads(blob)
|
||||
except ValueError:
|
||||
return None
|
||||
return data if isinstance(data, dict) else None
|
||||
|
||||
|
||||
def _decode_assignment(html: str, anchor: str) -> dict | None:
|
||||
"""Decode the balanced JSON object assigned after ``anchor``, or ``None``."""
|
||||
import json
|
||||
|
||||
idx = html.find(anchor)
|
||||
if idx == -1:
|
||||
return None
|
||||
brace = html.find("{", idx + len(anchor))
|
||||
if brace == -1:
|
||||
return None
|
||||
blob = _brace_match(html, brace)
|
||||
if not blob:
|
||||
return None
|
||||
try:
|
||||
data = json.loads(blob)
|
||||
except ValueError:
|
||||
return None
|
||||
return data if isinstance(data, dict) else None
|
||||
|
||||
|
||||
def extract_initial_data(html: str) -> dict | None:
|
||||
"""Return a /viewjob posting model rooted at ``jobInfoWrapperModel``.
|
||||
|
||||
Prefers ``window._rootProps`` (JSON) unwrapped at ``preloadedVJData``; falls
|
||||
back to a legacy inline ``window._initialData`` blob. ``window._initialData``
|
||||
is now a JS object literal that references other globals, so it is not JSON
|
||||
and is skipped when the JSON parse fails.
|
||||
"""
|
||||
root = _decode_assignment(html, _ROOT_PROPS_ANCHOR)
|
||||
if isinstance(root, dict):
|
||||
vj = root.get(_ROOT_PROPS_KEY)
|
||||
if isinstance(vj, dict) and vj.get("jobInfoWrapperModel"):
|
||||
return vj
|
||||
legacy = _decode_assignment(html, _INITIAL_DATA_ANCHOR)
|
||||
if isinstance(legacy, dict) and legacy.get("jobInfoWrapperModel"):
|
||||
return legacy
|
||||
return None
|
||||
|
||||
|
||||
def job_results(blob: dict | None) -> list[dict[str, Any]]:
|
||||
"""Return the raw job records from a decoded blob."""
|
||||
if not isinstance(blob, dict):
|
||||
return []
|
||||
results = (
|
||||
blob.get("metaData", {}).get("mosaicProviderJobCardsModel", {}).get("results")
|
||||
)
|
||||
if not isinstance(results, list):
|
||||
return []
|
||||
return [r for r in results if isinstance(r, dict)]
|
||||
|
||||
|
||||
def _utc_from_ms(value: Any) -> str | None:
|
||||
"""Epoch milliseconds -> millisecond ISO string."""
|
||||
if isinstance(value, bool) or not isinstance(value, int | float):
|
||||
return None
|
||||
dt = datetime.fromtimestamp(float(value) / 1000.0, tz=UTC)
|
||||
return dt.strftime("%Y-%m-%dT%H:%M:%S.%f")[:-3] + "Z"
|
||||
|
||||
|
||||
def _int(value: Any) -> int | None:
|
||||
"""Coerce to int, dropping bools."""
|
||||
if isinstance(value, bool):
|
||||
return None
|
||||
if isinstance(value, int):
|
||||
return value
|
||||
if isinstance(value, float):
|
||||
return int(value)
|
||||
return None
|
||||
|
||||
|
||||
def _abs_url(path: Any, base_url: str) -> str | None:
|
||||
"""Resolve an Indeed-relative path against ``base_url``; keep absolute URLs."""
|
||||
if not isinstance(path, str) or not path:
|
||||
return None
|
||||
if path.startswith("http"):
|
||||
return path
|
||||
return f"{base_url}{path if path.startswith('/') else '/' + path}"
|
||||
|
||||
|
||||
def _clean_snippet(snippet: Any) -> str | None:
|
||||
"""Strip tags and decode entities into plain text."""
|
||||
if not isinstance(snippet, str) or not snippet:
|
||||
return None
|
||||
text = _re_sub(r"<[^>]+>", " ", snippet)
|
||||
text = unescape(text)
|
||||
return _re_sub(r"\s+", " ", text).strip() or None
|
||||
|
||||
|
||||
def _taxonomy(raw: dict[str, Any]) -> dict[str, list[str]]:
|
||||
"""Flatten ``taxonomyAttributes`` into ``{group label: [attribute labels]}``."""
|
||||
out: dict[str, list[str]] = {}
|
||||
for group in raw.get("taxonomyAttributes") or []:
|
||||
if not isinstance(group, dict):
|
||||
continue
|
||||
label = group.get("label")
|
||||
attrs = group.get("attributes")
|
||||
if isinstance(label, str) and isinstance(attrs, list):
|
||||
out[label] = [
|
||||
a["label"]
|
||||
for a in attrs
|
||||
if isinstance(a, dict) and isinstance(a.get("label"), str)
|
||||
]
|
||||
return out
|
||||
|
||||
|
||||
def _job_types(raw: dict[str, Any], taxo: dict[str, list[str]]) -> list[str]:
|
||||
"""Job types from ``jobTypes`` then the taxonomy, deduped and order-stable."""
|
||||
seen: dict[str, None] = {}
|
||||
for jt in raw.get("jobTypes") or []:
|
||||
if isinstance(jt, str):
|
||||
seen.setdefault(jt, None)
|
||||
for label in ("job-types", "job-types-cc"):
|
||||
for jt in taxo.get(label, []):
|
||||
seen.setdefault(jt, None)
|
||||
return list(seen)
|
||||
|
||||
|
||||
def _salary(raw: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Flatten salary from ``salarySnippet`` (text) + ``extractedSalary`` (bounds)."""
|
||||
snippet = raw.get("salarySnippet") or {}
|
||||
extracted = raw.get("extractedSalary") or {}
|
||||
estimated = raw.get("estimatedSalary") or {}
|
||||
source = extracted or estimated
|
||||
text = snippet.get("text") if isinstance(snippet, dict) else None
|
||||
return {
|
||||
"salaryText": text if isinstance(text, str) else None,
|
||||
"salaryMin": source.get("min") if isinstance(source, dict) else None,
|
||||
"salaryMax": source.get("max") if isinstance(source, dict) else None,
|
||||
"currency": snippet.get("currency") if isinstance(snippet, dict) else None,
|
||||
"period": _SALARY_PERIODS.get(
|
||||
source.get("type") if isinstance(source, dict) else None
|
||||
),
|
||||
"isEstimated": bool(estimated) and not extracted,
|
||||
}
|
||||
|
||||
|
||||
def _is_remote(raw: dict[str, Any], taxo: dict[str, list[str]]) -> bool:
|
||||
"""Resolve remote/hybrid across Indeed's several signals."""
|
||||
if raw.get("remoteLocation") is True:
|
||||
return True
|
||||
if isinstance(raw.get("remoteWorkModel"), dict):
|
||||
return True
|
||||
return bool(taxo.get("remote"))
|
||||
|
||||
|
||||
def parse_job(raw: dict[str, Any], *, base_url: str = _DEFAULT_BASE) -> dict[str, Any]:
|
||||
"""Map one raw ``results[]`` record to a flat item dict.
|
||||
|
||||
``base_url`` is the country domain the record came from, so job and company
|
||||
URLs resolve to the right host.
|
||||
"""
|
||||
taxo = _taxonomy(raw)
|
||||
job_key = raw.get("jobkey")
|
||||
remote_model = raw.get("remoteWorkModel")
|
||||
return {
|
||||
"jobKey": job_key if isinstance(job_key, str) else None,
|
||||
"title": raw.get("displayTitle") or raw.get("title"),
|
||||
"jobUrl": f"{base_url}/viewjob?jk={job_key}" if job_key else None,
|
||||
"applyUrl": raw.get("thirdPartyApplyUrl") or None,
|
||||
"company": raw.get("company") or raw.get("truncatedCompany"),
|
||||
"companyUrl": _abs_url(raw.get("companyOverviewLink"), base_url),
|
||||
"companyRating": raw.get("companyRating"),
|
||||
"companyReviewCount": _int(raw.get("companyReviewCount")),
|
||||
"formattedLocation": raw.get("formattedLocation"),
|
||||
"city": raw.get("jobLocationCity"),
|
||||
"state": raw.get("jobLocationState"),
|
||||
"postalCode": raw.get("jobLocationPostal"),
|
||||
"country": raw.get("country"),
|
||||
"isRemote": _is_remote(raw, taxo),
|
||||
"remoteType": remote_model.get("type")
|
||||
if isinstance(remote_model, dict)
|
||||
else None,
|
||||
"jobTypes": _job_types(raw, taxo),
|
||||
"salary": _salary(raw),
|
||||
"benefits": taxo.get("benefits", []),
|
||||
"descriptionText": _clean_snippet(raw.get("snippet")),
|
||||
"descriptionHtml": None,
|
||||
"sponsored": raw.get("sponsored"),
|
||||
"isNew": raw.get("newJob"),
|
||||
"urgentlyHiring": raw.get("urgentlyHiring"),
|
||||
"expired": raw.get("expired"),
|
||||
"indeedApplyEnabled": raw.get("indeedApplyEnabled"),
|
||||
"age": raw.get("formattedRelativeTime"),
|
||||
"datePublished": _utc_from_ms(raw.get("pubDate")),
|
||||
"createdAt": _utc_from_ms(raw.get("createDate")),
|
||||
}
|
||||
|
||||
|
||||
def _detail_salary(hdr: dict[str, Any]) -> dict[str, Any] | None:
|
||||
"""Salary from the detail header's flat ``salaryMin/Max/Currency/Type`` fields."""
|
||||
smin = hdr.get("salaryMin")
|
||||
smax = hdr.get("salaryMax")
|
||||
if smin is None and smax is None:
|
||||
return None
|
||||
return {
|
||||
"salaryText": None,
|
||||
"salaryMin": smin,
|
||||
"salaryMax": smax,
|
||||
"currency": hdr.get("salaryCurrency"),
|
||||
"period": _SALARY_PERIODS.get(hdr.get("salaryType")),
|
||||
"isEstimated": False,
|
||||
}
|
||||
|
||||
|
||||
def parse_job_detail(html: str, *, base_url: str = _DEFAULT_BASE) -> dict[str, Any]:
|
||||
"""Map a /viewjob page to enrichment fields (empty dict if not a job page).
|
||||
|
||||
Returns only fields the detail page actually carries, so the caller can merge
|
||||
it onto a listing item without clobbering known values with blanks. The full
|
||||
description (``sanitizedJobDescription``) is the field listings never have.
|
||||
"""
|
||||
data = extract_initial_data(html)
|
||||
if not isinstance(data, dict):
|
||||
return {}
|
||||
jim = (data.get("jobInfoWrapperModel") or {}).get("jobInfoModel") or {}
|
||||
if not isinstance(jim, dict):
|
||||
return {}
|
||||
hdr = jim.get("jobInfoHeaderModel")
|
||||
hdr = hdr if isinstance(hdr, dict) else {}
|
||||
taxo = _taxonomy(hdr)
|
||||
desc_html = jim.get("sanitizedJobDescription")
|
||||
desc_html = desc_html if isinstance(desc_html, str) and desc_html else None
|
||||
remote_model = hdr.get("remoteWorkModel")
|
||||
out: dict[str, Any] = {
|
||||
"descriptionHtml": desc_html,
|
||||
"descriptionText": _clean_snippet(desc_html),
|
||||
"title": hdr.get("jobTitle"),
|
||||
"company": hdr.get("companyName"),
|
||||
"companyUrl": _abs_url(hdr.get("companyOverviewLink"), base_url),
|
||||
"formattedLocation": hdr.get("formattedLocation") or data.get("jobLocation"),
|
||||
"remoteType": remote_model.get("type")
|
||||
if isinstance(remote_model, dict)
|
||||
else None,
|
||||
"jobTypes": _job_types(hdr, taxo),
|
||||
"benefits": taxo.get("benefits", []),
|
||||
"salary": _detail_salary(hdr),
|
||||
}
|
||||
if _is_remote(hdr, taxo):
|
||||
out["isRemote"] = True
|
||||
return {k: v for k, v in out.items() if v not in (None, [], {})}
|
||||
|
|
@ -0,0 +1,122 @@
|
|||
# ruff: noqa: N815 - field names intentionally use the public camelCase API
|
||||
"""Input/output models for the Indeed scraper.
|
||||
|
||||
Anonymous scraper: there is no auth field. Fields absent from a listing (full
|
||||
description, benefits) stay ``None``/``[]`` until a detail fetch fills them.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Any, Literal
|
||||
|
||||
from pydantic import BaseModel, ConfigDict, Field
|
||||
|
||||
IndeedSort = Literal["relevance", "date"]
|
||||
IndeedJobType = Literal[
|
||||
"fulltime",
|
||||
"parttime",
|
||||
"contract",
|
||||
"internship",
|
||||
"temporary",
|
||||
"permanent",
|
||||
"seasonal",
|
||||
"freelance",
|
||||
]
|
||||
IndeedLevel = Literal["entry_level", "mid_level", "senior_level"]
|
||||
IndeedRemote = Literal["remote", "hybrid"]
|
||||
SalaryPeriod = Literal["hour", "day", "week", "month", "year"]
|
||||
|
||||
|
||||
class StartUrl(BaseModel):
|
||||
"""A direct URL entry; extra keys ignored."""
|
||||
|
||||
model_config = ConfigDict(extra="allow")
|
||||
|
||||
url: str
|
||||
|
||||
|
||||
class IndeedScrapeInput(BaseModel):
|
||||
"""Indeed scraper input. Caps are collector policy, enforced by ``scrape_indeed``."""
|
||||
|
||||
model_config = ConfigDict(extra="allow")
|
||||
|
||||
# Discovery: direct URLs and/or search queries.
|
||||
startUrls: list[StartUrl] = Field(default_factory=list)
|
||||
queries: list[str] = Field(default_factory=list)
|
||||
|
||||
# Search parameters applied to ``queries``.
|
||||
country: str = "us"
|
||||
location: str | None = None
|
||||
radius: int | None = None
|
||||
jobType: IndeedJobType | None = None
|
||||
level: IndeedLevel | None = None
|
||||
remote: IndeedRemote | None = None
|
||||
fromDays: int | None = None
|
||||
sort: IndeedSort = "relevance"
|
||||
|
||||
# Fetch each job's detail page for the full description.
|
||||
scrapeJobDetails: bool = False
|
||||
|
||||
maxItems: int = Field(default=25, ge=0)
|
||||
maxItemsPerQuery: int = Field(default=25, ge=0)
|
||||
|
||||
|
||||
class Salary(BaseModel):
|
||||
"""Salary block; fields are ``None`` when Indeed omits pay."""
|
||||
|
||||
model_config = ConfigDict(extra="allow")
|
||||
|
||||
salaryText: str | None = None
|
||||
salaryMin: float | None = None
|
||||
salaryMax: float | None = None
|
||||
currency: str | None = None
|
||||
period: SalaryPeriod | None = None
|
||||
isEstimated: bool | None = None
|
||||
|
||||
|
||||
class IndeedItem(BaseModel):
|
||||
"""One job posting. ``extra="allow"`` keeps the contract additive."""
|
||||
|
||||
model_config = ConfigDict(extra="allow")
|
||||
|
||||
jobKey: str | None = None
|
||||
title: str | None = None
|
||||
jobUrl: str | None = None
|
||||
applyUrl: str | None = None
|
||||
|
||||
company: str | None = None
|
||||
companyUrl: str | None = None
|
||||
companyRating: float | None = None
|
||||
companyReviewCount: int | None = None
|
||||
|
||||
formattedLocation: str | None = None
|
||||
city: str | None = None
|
||||
state: str | None = None
|
||||
postalCode: str | None = None
|
||||
country: str | None = None
|
||||
isRemote: bool | None = None
|
||||
remoteType: str | None = None
|
||||
|
||||
jobTypes: list[str] = Field(default_factory=list)
|
||||
salary: Salary = Field(default_factory=Salary)
|
||||
benefits: list[str] = Field(default_factory=list)
|
||||
|
||||
descriptionText: str | None = None
|
||||
descriptionHtml: str | None = None
|
||||
|
||||
sponsored: bool | None = None
|
||||
isNew: bool | None = None
|
||||
urgentlyHiring: bool | None = None
|
||||
expired: bool | None = None
|
||||
indeedApplyEnabled: bool | None = None
|
||||
|
||||
age: str | None = None
|
||||
datePublished: str | None = None
|
||||
createdAt: str | None = None
|
||||
scrapedAt: str | None = None
|
||||
|
||||
source: str = "indeed"
|
||||
|
||||
def to_output(self) -> dict[str, Any]:
|
||||
"""Serialize to the flat output dict, keeping extras."""
|
||||
return self.model_dump(exclude_none=False)
|
||||
|
|
@ -0,0 +1,178 @@
|
|||
"""Orchestrator for the Indeed scraper.
|
||||
|
||||
:func:`iter_indeed` streams deduped job items from one warmed session; each
|
||||
search/company target contributes its first page. :func:`scrape_indeed` collects
|
||||
the stream under a caller ``limit``. Targets run sequentially to reuse the session.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import logging
|
||||
from collections.abc import AsyncIterator
|
||||
from typing import Any
|
||||
from urllib.parse import urlparse
|
||||
|
||||
from .fetch import IndeedSession, now_iso, open_session
|
||||
from .parsers import (
|
||||
extract_jobcards_blob,
|
||||
job_results,
|
||||
parse_job,
|
||||
parse_job_detail,
|
||||
)
|
||||
from .schemas import IndeedItem, IndeedScrapeInput
|
||||
from .url_resolver import build_search_url, resolve_url
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
|
||||
__all__ = ["iter_indeed", "scrape_indeed"]
|
||||
|
||||
|
||||
def _emit(partial: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Stamp ``scrapedAt`` and normalize through the output model."""
|
||||
return IndeedItem(**{**partial, "scrapedAt": now_iso()}).to_output()
|
||||
|
||||
|
||||
async def _search_items(
|
||||
session: IndeedSession, url: str, *, domain: str, max_items: int
|
||||
) -> AsyncIterator[dict[str, Any]]:
|
||||
"""Yield deduped job cards from one search/company page.
|
||||
|
||||
ponytail: caps a query at its first page (~15 jobs) — anonymous Indeed gates
|
||||
``start>=10``; deeper depth needs an authenticated session or Indeed's API.
|
||||
"""
|
||||
if max_items <= 0:
|
||||
return
|
||||
base_url = f"https://{domain}"
|
||||
html = await session.fetch_html(url)
|
||||
seen: set[str] = set()
|
||||
emitted = 0
|
||||
for raw in job_results(extract_jobcards_blob(html)):
|
||||
item = parse_job(raw, base_url=base_url)
|
||||
job_key = item.get("jobKey")
|
||||
if isinstance(job_key, str):
|
||||
if job_key in seen:
|
||||
continue
|
||||
seen.add(job_key)
|
||||
yield _emit(item)
|
||||
emitted += 1
|
||||
if emitted >= max_items:
|
||||
return
|
||||
|
||||
|
||||
def _targets(input_model: IndeedScrapeInput) -> list[tuple[str, str, str]]:
|
||||
"""Resolve inputs to ``(kind, url, domain)`` targets.
|
||||
|
||||
``startUrls`` take precedence over ``queries``. ``kind`` is ``search`` for
|
||||
query-built and search/company URLs, or ``job`` for a ``/viewjob`` URL.
|
||||
"""
|
||||
if input_model.startUrls:
|
||||
out: list[tuple[str, str, str]] = []
|
||||
for entry in input_model.startUrls:
|
||||
resolved = resolve_url(entry.url)
|
||||
if resolved is None:
|
||||
logger.warning("[indeed] skipping unrecognized URL: %s", entry.url)
|
||||
continue
|
||||
kind = "job" if resolved.kind == "job" else "search"
|
||||
out.append((kind, resolved.url, resolved.domain))
|
||||
return out
|
||||
|
||||
domain = None
|
||||
urls: list[tuple[str, str, str]] = []
|
||||
for query in input_model.queries:
|
||||
url = build_search_url(
|
||||
query,
|
||||
country=input_model.country,
|
||||
location=input_model.location,
|
||||
radius=input_model.radius,
|
||||
job_type=input_model.jobType,
|
||||
level=input_model.level,
|
||||
remote=input_model.remote,
|
||||
from_days=input_model.fromDays,
|
||||
sort=input_model.sort,
|
||||
)
|
||||
domain = domain or urlparse(url).hostname or "www.indeed.com"
|
||||
urls.append(("search", url, domain))
|
||||
return urls
|
||||
|
||||
|
||||
async def _enrich(session: IndeedSession, item: dict[str, Any], base_url: str) -> None:
|
||||
"""Merge a job's /viewjob detail (full description, etc.) onto ``item`` in place.
|
||||
|
||||
Best-effort: a blocked or malformed detail page leaves the listing fields as-is
|
||||
rather than failing the run.
|
||||
"""
|
||||
job_url = item.get("jobUrl")
|
||||
if not isinstance(job_url, str):
|
||||
return
|
||||
try:
|
||||
# Fail fast: enrichment is best-effort, so a gated detail page must not
|
||||
# rotate IPs and eat the run's time budget for one job's description.
|
||||
html = await session.fetch_html(job_url, max_rotations=0)
|
||||
detail = parse_job_detail(html, base_url=base_url)
|
||||
except Exception as exc:
|
||||
logger.warning("[indeed] detail fetch failed for %s: %s", job_url, exc)
|
||||
return
|
||||
item.update(detail)
|
||||
|
||||
|
||||
async def _job_item(
|
||||
session: IndeedSession, url: str, base_url: str
|
||||
) -> dict[str, Any] | None:
|
||||
"""Scrape a single /viewjob URL into an item from its detail page alone."""
|
||||
detail = parse_job_detail(await session.fetch_html(url), base_url=base_url)
|
||||
if not detail:
|
||||
return None
|
||||
return _emit({"jobUrl": url, "source": "indeed", **detail})
|
||||
|
||||
|
||||
async def iter_indeed(
|
||||
input_model: IndeedScrapeInput, session: IndeedSession
|
||||
) -> AsyncIterator[dict[str, Any]]:
|
||||
"""Stream flat job items for every target, deduped by ``jobKey`` across all."""
|
||||
global_seen: set[str] = set()
|
||||
for kind, url, domain in _targets(input_model):
|
||||
base_url = f"https://{domain}"
|
||||
if kind == "job":
|
||||
item = await _job_item(session, url, base_url)
|
||||
if item is not None:
|
||||
yield item
|
||||
continue
|
||||
async for item in _search_items(
|
||||
session, url, domain=domain, max_items=input_model.maxItemsPerQuery
|
||||
):
|
||||
job_key = item.get("jobKey")
|
||||
if isinstance(job_key, str):
|
||||
if job_key in global_seen:
|
||||
continue
|
||||
global_seen.add(job_key)
|
||||
if input_model.scrapeJobDetails:
|
||||
await _enrich(session, item, base_url)
|
||||
yield item
|
||||
|
||||
|
||||
async def scrape_indeed(
|
||||
input_model: IndeedScrapeInput,
|
||||
*,
|
||||
limit: int | None = None,
|
||||
session: IndeedSession | None = None,
|
||||
) -> list[dict[str, Any]]:
|
||||
"""Collect :func:`iter_indeed` into a list under an optional ``limit``.
|
||||
|
||||
Opens a warmed session when one is not supplied. ``limit`` is a request-time
|
||||
guard, not a ceiling baked into the stream.
|
||||
"""
|
||||
from app.capabilities.core.progress import emit_progress
|
||||
|
||||
async def _collect(sess: IndeedSession) -> list[dict[str, Any]]:
|
||||
results: list[dict[str, Any]] = []
|
||||
async for item in iter_indeed(input_model, sess):
|
||||
results.append(item)
|
||||
emit_progress("scraping", current=len(results), total=limit, unit="item")
|
||||
if limit is not None and len(results) >= limit:
|
||||
break
|
||||
return results
|
||||
|
||||
if session is not None:
|
||||
return await _collect(session)
|
||||
async with open_session() as sess:
|
||||
return await _collect(sess)
|
||||
|
|
@ -0,0 +1,124 @@
|
|||
"""Classify Indeed URLs and build search URLs.
|
||||
|
||||
Recognizes search pages (``/jobs?q=&l=``), company pages (``/cmp/<slug>/jobs``),
|
||||
and single jobs (``/viewjob?jk=``); other hosts resolve to ``None``. Also owns
|
||||
the country->domain map so classification and URL building share one source.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass, field
|
||||
from typing import Literal
|
||||
from urllib.parse import parse_qs, urlencode, urlparse
|
||||
|
||||
ResolvedKind = Literal["search", "company", "job"]
|
||||
|
||||
# Locale subdomains that deviate from the ISO code; others map to <cc>.indeed.com.
|
||||
_DOMAIN_OVERRIDES = {"us": "www", "gb": "uk"}
|
||||
|
||||
_JT_VALUES = frozenset(
|
||||
{
|
||||
"fulltime",
|
||||
"parttime",
|
||||
"contract",
|
||||
"internship",
|
||||
"temporary",
|
||||
"permanent",
|
||||
"seasonal",
|
||||
"freelance",
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ResolvedUrl:
|
||||
kind: ResolvedKind
|
||||
value: str # search query, company slug, or job key
|
||||
url: str
|
||||
domain: str
|
||||
location: str | None = None
|
||||
params: dict[str, str] = field(default_factory=dict)
|
||||
|
||||
|
||||
def _is_indeed_host(hostname: str | None) -> bool:
|
||||
if not hostname:
|
||||
return False
|
||||
h = hostname.lower()
|
||||
return h == "indeed.com" or h.endswith(".indeed.com")
|
||||
|
||||
|
||||
def country_domain(country: str) -> str:
|
||||
"""Country code -> Indeed host, e.g. ``us`` -> ``www.indeed.com``."""
|
||||
cc = (country or "us").strip().lower()
|
||||
return f"{_DOMAIN_OVERRIDES.get(cc, cc)}.indeed.com"
|
||||
|
||||
|
||||
def resolve_url(url: str) -> ResolvedUrl | None:
|
||||
"""Classify an Indeed URL into a scrape job, or ``None`` if unrecognized."""
|
||||
parsed = urlparse(url)
|
||||
if not _is_indeed_host(parsed.hostname):
|
||||
return None
|
||||
domain = parsed.hostname or "www.indeed.com"
|
||||
path = (parsed.path or "").rstrip("/")
|
||||
query = parse_qs(parsed.query)
|
||||
segments = [s for s in path.split("/") if s]
|
||||
|
||||
# /viewjob?jk=<key>
|
||||
if path.endswith("/viewjob") or segments[:1] == ["viewjob"]:
|
||||
jk = query.get("jk", [None])[0]
|
||||
return ResolvedUrl("job", jk, url, domain) if jk else None
|
||||
|
||||
# /cmp/<slug>/jobs
|
||||
if segments[:1] == ["cmp"] and "jobs" in segments and len(segments) >= 2:
|
||||
return ResolvedUrl("company", segments[1], url, domain)
|
||||
|
||||
# /jobs?q=&l=
|
||||
if path.endswith("/jobs") or segments[-1:] == ["jobs"]:
|
||||
q = query.get("q", [""])[0]
|
||||
loc = query.get("l", [None])[0]
|
||||
extra = {
|
||||
k: v[0]
|
||||
for k, v in query.items()
|
||||
if k in ("radius", "sort", "fromage", "jt", "explvl") and v
|
||||
}
|
||||
return ResolvedUrl("search", q, url, domain, location=loc, params=extra)
|
||||
|
||||
return None
|
||||
|
||||
|
||||
def build_search_url(
|
||||
query: str,
|
||||
*,
|
||||
country: str = "us",
|
||||
location: str | None = None,
|
||||
radius: int | None = None,
|
||||
job_type: str | None = None,
|
||||
level: str | None = None,
|
||||
remote: str | None = None,
|
||||
from_days: int | None = None,
|
||||
sort: str = "relevance",
|
||||
start: int = 0,
|
||||
) -> str:
|
||||
"""Build an Indeed ``/jobs`` search URL.
|
||||
|
||||
Remote/hybrid is passed as a query keyword; Indeed's structured ``sc``
|
||||
attribute codes rotate and aren't stable to hardcode.
|
||||
"""
|
||||
domain = country_domain(country)
|
||||
q = f"{query} {remote}".strip() if remote else query
|
||||
params: dict[str, str] = {"q": q}
|
||||
if location:
|
||||
params["l"] = location
|
||||
if radius is not None:
|
||||
params["radius"] = str(radius)
|
||||
if job_type in _JT_VALUES:
|
||||
params["jt"] = job_type # type: ignore[assignment]
|
||||
if level:
|
||||
params["explvl"] = level
|
||||
if from_days is not None:
|
||||
params["fromage"] = str(from_days)
|
||||
if sort == "date":
|
||||
params["sort"] = "date"
|
||||
if start:
|
||||
params["start"] = str(start)
|
||||
return f"https://{domain}/jobs?{urlencode(params)}"
|
||||
|
|
@ -4,6 +4,7 @@ from fastapi import APIRouter, Depends
|
|||
import app.capabilities.amazon
|
||||
import app.capabilities.google_maps
|
||||
import app.capabilities.google_search
|
||||
import app.capabilities.indeed
|
||||
import app.capabilities.instagram
|
||||
import app.capabilities.reddit
|
||||
import app.capabilities.tiktok
|
||||
|
|
|
|||
149
surfsense_backend/scripts/e2e_indeed_scraper.py
Normal file
149
surfsense_backend/scripts/e2e_indeed_scraper.py
Normal file
|
|
@ -0,0 +1,149 @@
|
|||
"""Manual functional e2e for the Indeed scraper (app/proprietary/platforms/indeed_jobs).
|
||||
|
||||
Run from the backend directory:
|
||||
cd surfsense_backend
|
||||
uv run python scripts/e2e_indeed_scraper.py
|
||||
# or: .venv/bin/python scripts/e2e_indeed_scraper.py
|
||||
|
||||
This is NOT a pytest test (it needs live network + a residential/custom proxy).
|
||||
All steps share one warmed browser session:
|
||||
|
||||
Step 0 — go/no-go probe: open a session, fetch a live search page, and assert
|
||||
its embedded job-cards blob parses into results. If this fails the whole
|
||||
approach is blocked on this IP/proxy — later steps are skipped.
|
||||
Step 1 — search query -> job items; keep one discovered /viewjob URL.
|
||||
Step 2 — scrape a search URL via startUrls.
|
||||
Step 3 — scrape the discovered /viewjob URL and assert a full description.
|
||||
Step 4 — scrape_job_details enrichment: a query with detail pages fetched.
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
from dotenv import load_dotenv
|
||||
|
||||
# --- bootstrap: load .env and put the backend root on sys.path before app.* ---
|
||||
_BACKEND_ROOT = Path(__file__).resolve().parent.parent
|
||||
sys.path.insert(0, str(_BACKEND_ROOT))
|
||||
for _candidate in (_BACKEND_ROOT / ".env", _BACKEND_ROOT.parent / ".env"):
|
||||
if _candidate.exists():
|
||||
load_dotenv(_candidate)
|
||||
break
|
||||
|
||||
from app.proprietary.platforms.indeed_jobs import ( # noqa: E402
|
||||
IndeedScrapeInput,
|
||||
scrape_indeed,
|
||||
)
|
||||
from app.proprietary.platforms.indeed_jobs.fetch import open_session # noqa: E402
|
||||
from app.proprietary.platforms.indeed_jobs.parsers import ( # noqa: E402
|
||||
extract_jobcards_blob,
|
||||
job_results,
|
||||
)
|
||||
from app.proprietary.platforms.indeed_jobs.url_resolver import ( # noqa: E402
|
||||
build_search_url,
|
||||
)
|
||||
|
||||
_QUERY = "data analyst"
|
||||
_LOCATION = "Remote"
|
||||
|
||||
|
||||
def _hr(title: str) -> None:
|
||||
print(f"\n{'=' * 70}\n{title}\n{'=' * 70}")
|
||||
|
||||
|
||||
def _check(label: str, ok: bool, detail: str = "") -> bool:
|
||||
print(f" [{'PASS' if ok else 'FAIL'}] {label}{f' — {detail}' if detail else ''}")
|
||||
return ok
|
||||
|
||||
|
||||
async def step0_probe(sess, state: dict) -> bool:
|
||||
_hr("STEP 0 — go/no-go: warmed session + parseable job cards")
|
||||
url = build_search_url(_QUERY, location=_LOCATION)
|
||||
html = await sess.fetch_html(url)
|
||||
raws = job_results(extract_jobcards_blob(html))
|
||||
print(f" {url} -> job_cards={len(raws)}")
|
||||
return _check("search page parsed job cards", len(raws) > 0, f"{len(raws)} cards")
|
||||
|
||||
|
||||
async def step1_search(sess, state: dict) -> bool:
|
||||
_hr("STEP 1 — search query -> items")
|
||||
items = await scrape_indeed(
|
||||
IndeedScrapeInput(queries=[_QUERY], location=_LOCATION, maxItems=5),
|
||||
limit=5,
|
||||
session=sess,
|
||||
)
|
||||
for it in items[:5]:
|
||||
print(f" - {it.get('jobKey')} | {it.get('title')} @ {it.get('company')}")
|
||||
state["job_url"] = next(
|
||||
(it["jobUrl"] for it in items if it.get("jobUrl")), None
|
||||
)
|
||||
return _check("search returned jobs", len(items) > 0, f"{len(items)} jobs")
|
||||
|
||||
|
||||
async def step2_search_url(sess, state: dict) -> bool:
|
||||
_hr("STEP 2 — scrape a search URL (startUrls)")
|
||||
url = build_search_url(_QUERY, location=_LOCATION)
|
||||
items = await scrape_indeed(
|
||||
IndeedScrapeInput(startUrls=[{"url": url}], maxItems=5),
|
||||
limit=5,
|
||||
session=sess,
|
||||
)
|
||||
return _check("search URL returned jobs", len(items) > 0, f"{len(items)} jobs")
|
||||
|
||||
|
||||
async def step3_viewjob(sess, state: dict) -> bool:
|
||||
_hr("STEP 3 — scrape a discovered /viewjob URL (full description)")
|
||||
url = state.get("job_url")
|
||||
if not url:
|
||||
return _check("had a discovered job URL", False)
|
||||
items = await scrape_indeed(
|
||||
IndeedScrapeInput(startUrls=[{"url": url}], maxItems=1),
|
||||
limit=1,
|
||||
session=sess,
|
||||
)
|
||||
desc = items[0].get("descriptionText") if items else None
|
||||
print(f" url={url}\n description_chars={len(desc or '')}")
|
||||
return _check("viewjob returned a description", bool(desc), url)
|
||||
|
||||
|
||||
async def step4_enrich(sess, state: dict) -> bool:
|
||||
_hr("STEP 4 — search with scrape_job_details enrichment")
|
||||
items = await scrape_indeed(
|
||||
IndeedScrapeInput(
|
||||
queries=[_QUERY],
|
||||
location=_LOCATION,
|
||||
maxItems=3,
|
||||
scrapeJobDetails=True,
|
||||
),
|
||||
limit=3,
|
||||
session=sess,
|
||||
)
|
||||
# descriptionHtml is set only by the detail page; listings never carry it,
|
||||
# so it proves enrichment actually merged the /viewjob model.
|
||||
enriched = [i for i in items if i.get("descriptionHtml")]
|
||||
return _check(
|
||||
"enriched items carry full descriptionHtml",
|
||||
bool(enriched),
|
||||
f"{len(enriched)}/{len(items)} enriched",
|
||||
)
|
||||
|
||||
|
||||
async def main() -> int:
|
||||
state: dict = {}
|
||||
async with open_session() as sess:
|
||||
results = [await step0_probe(sess, state)]
|
||||
if not results[-1]:
|
||||
print("\nprobe failed — Indeed is blocking this IP/proxy. Aborting.")
|
||||
return 1
|
||||
results.append(await step1_search(sess, state))
|
||||
results.append(await step2_search_url(sess, state))
|
||||
results.append(await step3_viewjob(sess, state))
|
||||
results.append(await step4_enrich(sess, state))
|
||||
_hr("SUMMARY")
|
||||
print(f" {sum(results)}/{len(results)} steps passed")
|
||||
return 0 if all(results) else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(asyncio.run(main()))
|
||||
|
|
@ -33,6 +33,7 @@ _EXPECTED_SUBAGENTS = frozenset(
|
|||
"google_drive",
|
||||
"google_maps",
|
||||
"google_search",
|
||||
"indeed",
|
||||
"instagram",
|
||||
"knowledge_base",
|
||||
"mcp_discovery",
|
||||
|
|
|
|||
|
|
@ -0,0 +1,100 @@
|
|||
"""``indeed.scrape`` executor: verb input → actor input mapping → typed items.
|
||||
|
||||
Boundary mocked: the proprietary scraper (injected fake). NOT mocked: the verb's
|
||||
own payload→IndeedScrapeInput mapping and the dict→IndeedItem wrapping.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from app.capabilities.indeed.scrape.executor import build_scrape_executor
|
||||
from app.capabilities.indeed.scrape.schemas import ScrapeInput, ScrapeOutput
|
||||
from app.exceptions import ForbiddenError
|
||||
from app.proprietary.platforms.indeed_jobs import (
|
||||
IndeedAccessBlockedError,
|
||||
IndeedScrapeInput,
|
||||
)
|
||||
|
||||
pytestmark = pytest.mark.unit
|
||||
|
||||
|
||||
class _FakeScraper:
|
||||
"""Records the actor input + limit it was called with; returns canned items."""
|
||||
|
||||
def __init__(self, items: list[dict]):
|
||||
self._items = items
|
||||
self.calls: list[tuple[IndeedScrapeInput, int | None]] = []
|
||||
|
||||
async def __call__(
|
||||
self, actor_input: IndeedScrapeInput, *, limit: int | None = None
|
||||
) -> list[dict]:
|
||||
self.calls.append((actor_input, limit))
|
||||
return self._items
|
||||
|
||||
|
||||
async def test_maps_urls_to_start_urls_and_wraps_items():
|
||||
scraper = _FakeScraper([{"jobKey": "abc", "title": "Data Analyst"}])
|
||||
execute = build_scrape_executor(scrape_fn=scraper)
|
||||
|
||||
out = await execute(ScrapeInput(urls=["https://www.indeed.com/jobs?q=dev"]))
|
||||
|
||||
assert isinstance(out, ScrapeOutput)
|
||||
assert len(out.items) == 1
|
||||
assert out.items[0].jobKey == "abc"
|
||||
assert out.items[0].title == "Data Analyst"
|
||||
|
||||
(actor_input, _limit) = scraper.calls[0]
|
||||
assert [u.url for u in actor_input.startUrls] == [
|
||||
"https://www.indeed.com/jobs?q=dev"
|
||||
]
|
||||
assert actor_input.queries == []
|
||||
|
||||
|
||||
async def test_maps_search_params_and_passes_limit():
|
||||
scraper = _FakeScraper([])
|
||||
execute = build_scrape_executor(scrape_fn=scraper)
|
||||
|
||||
await execute(
|
||||
ScrapeInput(
|
||||
search_queries=["data analyst"],
|
||||
country="gb",
|
||||
location="Remote",
|
||||
job_type="fulltime",
|
||||
level="entry_level",
|
||||
remote="remote",
|
||||
from_days=7,
|
||||
sort="date",
|
||||
max_items=40,
|
||||
max_items_per_query=15,
|
||||
)
|
||||
)
|
||||
|
||||
(actor_input, limit) = scraper.calls[0]
|
||||
assert actor_input.queries == ["data analyst"]
|
||||
assert actor_input.country == "gb"
|
||||
assert actor_input.location == "Remote"
|
||||
assert actor_input.jobType == "fulltime"
|
||||
assert actor_input.level == "entry_level"
|
||||
assert actor_input.remote == "remote"
|
||||
assert actor_input.fromDays == 7
|
||||
assert actor_input.sort == "date"
|
||||
assert actor_input.maxItems == 40
|
||||
assert actor_input.maxItemsPerQuery == 15
|
||||
# The outer collection limit is the caller's total-item cap.
|
||||
assert limit == 40
|
||||
|
||||
|
||||
async def test_access_blocked_maps_to_forbidden():
|
||||
async def _blocked(actor_input: IndeedScrapeInput, *, limit: int | None = None):
|
||||
raise IndeedAccessBlockedError("all IPs refused")
|
||||
|
||||
execute = build_scrape_executor(scrape_fn=_blocked)
|
||||
|
||||
with pytest.raises(ForbiddenError):
|
||||
await execute(ScrapeInput(search_queries=["x"]))
|
||||
|
||||
|
||||
def test_requires_a_source():
|
||||
with pytest.raises(ValueError, match="at least one"):
|
||||
ScrapeInput()
|
||||
|
|
@ -0,0 +1,23 @@
|
|||
"""The indeed namespace registers its verb as one Capability the doors/agent read."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from app.capabilities import (
|
||||
indeed, # noqa: F401 — importing the namespace registers its verbs
|
||||
)
|
||||
from app.capabilities.core.store import get_capability
|
||||
from app.capabilities.core.types import BillingUnit
|
||||
from app.capabilities.indeed.scrape.schemas import ScrapeInput, ScrapeOutput
|
||||
|
||||
pytestmark = pytest.mark.unit
|
||||
|
||||
|
||||
def test_indeed_scrape_is_registered_and_billed_per_job():
|
||||
cap = get_capability("indeed.scrape")
|
||||
|
||||
assert cap.name == "indeed.scrape"
|
||||
assert cap.input_schema is ScrapeInput
|
||||
assert cap.output_schema is ScrapeOutput
|
||||
assert cap.billing_unit is BillingUnit.INDEED_JOB
|
||||
File diff suppressed because it is too large
Load diff
|
|
@ -0,0 +1 @@
|
|||
<!doctype html><html><head><script>window._initialData = {"jobLocation": "Southington, CT 06489", "jobInfoWrapperModel": {"jobInfoModel": {"sanitizedJobDescription": "<div>\n <div>\n <div>\n <b>Hybrid Data Analyst in Southington, CT \u2013 Contact Center Analytics & Conversational Intelligence (BPO)</b>\n </div>\n </div>\n <p><b> Overview</b></p>\n <p>We are seeking a Data Analyst to support a BPO contact center environment servicing multiple clients. The role focuses on transforming complex interaction and performance data into clear, actionable insights. Candidat</div>", "jobInfoHeaderModel": {"jobTitle": "Hybrid Data Analyst in Southington, CT", "companyName": "Company Confidential", "companyOverviewLink": "https://www.indeed.com/cmp/Company-Confidential-1?campaignid=mobvjcmp&from=mobviewjob&tk=1jth31e29ijvr800&fromjk=703094f9d01d4079", "formattedLocation": "Southington, CT 06489", "salaryMin": 60000, "salaryMax": 90000, "salaryCurrency": "USD", "salaryType": "YEARLY", "remoteWorkModel": {"type": "HYBRID"}, "jobTypes": null, "taxonomyAttributes": [{"label": "job-types", "attributes": [{"label": "Full-time"}]}, {"label": "benefits", "attributes": [{"label": "401(k)"}, {"label": "Health insurance"}]}, {"label": "remote", "attributes": [{"label": "Hybrid work"}]}]}}}};</script></head><body></body></html>
|
||||
|
|
@ -0,0 +1,105 @@
|
|||
"""Offline tests for the rotate-on-block fetch loop (no network, fake session)."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import pytest
|
||||
|
||||
from app.proprietary.platforms.indeed_jobs.fetch import (
|
||||
IndeedAccessBlockedError,
|
||||
IndeedSession,
|
||||
)
|
||||
|
||||
_OK_HTML = "<html>jobs listing</html>"
|
||||
_BLOCK_HTML = "<html>secure.indeed.com security check</html>"
|
||||
|
||||
|
||||
class _FakePage:
|
||||
def __init__(self, html: str, url: str) -> None:
|
||||
self.html_content = html
|
||||
self.url = url
|
||||
|
||||
|
||||
class _Controller:
|
||||
"""Shared state across sessions the factory hands out (survives rotation)."""
|
||||
|
||||
def __init__(self, target_outcomes: list[str]) -> None:
|
||||
self.target_outcomes = target_outcomes
|
||||
self.sessions_started = 0
|
||||
self.home_fetches: dict[str, int] = {}
|
||||
self.target_index = 0
|
||||
|
||||
def factory(self) -> _FakeSession:
|
||||
return _FakeSession(self)
|
||||
|
||||
|
||||
class _FakeSession:
|
||||
def __init__(self, ctrl: _Controller) -> None:
|
||||
self._ctrl = ctrl
|
||||
|
||||
async def start(self) -> None:
|
||||
self._ctrl.sessions_started += 1
|
||||
|
||||
async def close(self) -> None:
|
||||
pass
|
||||
|
||||
async def fetch(self, url: str, **_: object) -> _FakePage:
|
||||
if url.endswith("/") and "/jobs" not in url: # warm-up hit
|
||||
self._ctrl.home_fetches[url] = self._ctrl.home_fetches.get(url, 0) + 1
|
||||
return _FakePage("<html>home</html>", url)
|
||||
outcome = self._ctrl.target_outcomes[self._ctrl.target_index]
|
||||
self._ctrl.target_index += 1
|
||||
if outcome == "OK":
|
||||
return _FakePage(_OK_HTML, url)
|
||||
if outcome == "ERROR":
|
||||
raise RuntimeError("boom")
|
||||
return _FakePage(_BLOCK_HTML, "https://secure.indeed.com/auth")
|
||||
|
||||
|
||||
_URL = "https://www.indeed.com/jobs?q=dev"
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_rotates_past_a_block_then_succeeds():
|
||||
ctrl = _Controller(["BLOCK", "OK"])
|
||||
session = IndeedSession(ctrl.factory)
|
||||
html = await session.fetch_html(_URL)
|
||||
assert html == _OK_HTML
|
||||
assert session.rotations == 1
|
||||
assert ctrl.sessions_started == 2 # initial + one rotation
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_recovers_after_a_fetch_error():
|
||||
ctrl = _Controller(["ERROR", "OK"])
|
||||
session = IndeedSession(ctrl.factory)
|
||||
assert await session.fetch_html(_URL) == _OK_HTML
|
||||
assert session.rotations == 1
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_raises_after_exhausting_rotations():
|
||||
ctrl = _Controller(["BLOCK"] * 10)
|
||||
session = IndeedSession(ctrl.factory)
|
||||
with pytest.raises(IndeedAccessBlockedError):
|
||||
await session.fetch_html(_URL)
|
||||
assert session.rotations == 3
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_max_rotations_zero_fails_fast():
|
||||
# A gated page (pagination) must raise on the first block without rotating.
|
||||
ctrl = _Controller(["BLOCK", "OK"])
|
||||
session = IndeedSession(ctrl.factory)
|
||||
with pytest.raises(IndeedAccessBlockedError):
|
||||
await session.fetch_html(_URL, max_rotations=0)
|
||||
assert session.rotations == 0
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_warms_domain_once_without_rotation():
|
||||
ctrl = _Controller(["OK", "OK"])
|
||||
session = IndeedSession(ctrl.factory)
|
||||
await session.fetch_html(_URL)
|
||||
await session.fetch_html(_URL + "&start=10")
|
||||
assert ctrl.home_fetches["https://www.indeed.com/"] == 1
|
||||
await session.close()
|
||||
|
|
@ -0,0 +1,196 @@
|
|||
"""Offline parser tests: synthetic mapping plus a real captured blob."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
|
||||
from app.proprietary.platforms.indeed_jobs.parsers import (
|
||||
extract_jobcards_blob,
|
||||
job_results,
|
||||
parse_job,
|
||||
parse_job_detail,
|
||||
)
|
||||
|
||||
_FIXTURE_DIR = Path(__file__).parent / "fixtures"
|
||||
|
||||
|
||||
# --- synthetic mapping (always runs) ---------------------------------------
|
||||
|
||||
|
||||
def _raw_job() -> dict:
|
||||
return {
|
||||
"jobkey": "abc123",
|
||||
"displayTitle": "Senior Data Analyst",
|
||||
"title": "Senior Data Analyst (fallback)",
|
||||
"company": "Acme Corp",
|
||||
"truncatedCompany": "Acme",
|
||||
"companyOverviewLink": "/cmp/Acme-Corp",
|
||||
"companyRating": 4.1,
|
||||
"companyReviewCount": 320,
|
||||
"formattedLocation": "New York, NY",
|
||||
"jobLocationCity": "New York",
|
||||
"jobLocationState": "NY",
|
||||
"jobLocationPostal": "10001",
|
||||
"country": "US",
|
||||
"remoteLocation": False,
|
||||
"remoteWorkModel": {"type": "REMOTE_ALWAYS", "text": "Remote"},
|
||||
"jobTypes": [],
|
||||
"salarySnippet": {"currency": "USD", "text": "$90,000 - $120,000 a year"},
|
||||
"extractedSalary": {"min": 90000, "max": 120000, "type": "YEARLY"},
|
||||
"snippet": "<ul><li>5+ years <b>SQL</b> & Python</li></ul>",
|
||||
"sponsored": True,
|
||||
"newJob": False,
|
||||
"urgentlyHiring": True,
|
||||
"expired": False,
|
||||
"indeedApplyEnabled": True,
|
||||
"formattedRelativeTime": "3 days ago",
|
||||
"pubDate": 1_774_242_000_000,
|
||||
"createDate": 1_774_276_267_415,
|
||||
"thirdPartyApplyUrl": "https://ats.example.com/apply/abc123",
|
||||
"taxonomyAttributes": [
|
||||
{"label": "job-types", "attributes": [{"label": "Full-time", "suid": "x"}]},
|
||||
{"label": "remote", "attributes": [{"label": "Remote", "suid": "y"}]},
|
||||
{
|
||||
"label": "benefits",
|
||||
"attributes": [
|
||||
{"label": "Health insurance", "suid": "a"},
|
||||
{"label": "401(k)", "suid": "b"},
|
||||
],
|
||||
},
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
def test_parse_job_maps_core_fields():
|
||||
item = parse_job(_raw_job())
|
||||
assert item["jobKey"] == "abc123"
|
||||
assert item["title"] == "Senior Data Analyst" # displayTitle wins over title
|
||||
assert item["jobUrl"] == "https://www.indeed.com/viewjob?jk=abc123"
|
||||
assert item["applyUrl"] == "https://ats.example.com/apply/abc123"
|
||||
assert item["company"] == "Acme Corp"
|
||||
assert item["companyUrl"] == "https://www.indeed.com/cmp/Acme-Corp"
|
||||
assert item["companyReviewCount"] == 320
|
||||
assert item["city"] == "New York"
|
||||
assert item["isRemote"] is True
|
||||
assert item["remoteType"] == "REMOTE_ALWAYS"
|
||||
assert item["jobTypes"] == ["Full-time"]
|
||||
assert item["benefits"] == ["Health insurance", "401(k)"]
|
||||
assert item["sponsored"] is True
|
||||
assert item["urgentlyHiring"] is True
|
||||
|
||||
|
||||
def test_parse_job_salary_and_snippet():
|
||||
item = parse_job(_raw_job())
|
||||
sal = item["salary"]
|
||||
assert sal["salaryText"] == "$90,000 - $120,000 a year"
|
||||
assert sal["salaryMin"] == 90000
|
||||
assert sal["salaryMax"] == 120000
|
||||
assert sal["currency"] == "USD"
|
||||
assert sal["period"] == "year"
|
||||
assert sal["isEstimated"] is False
|
||||
# snippet HTML is stripped + entities decoded into plain text.
|
||||
assert item["descriptionText"] == "5+ years SQL & Python"
|
||||
assert item["descriptionHtml"] is None
|
||||
|
||||
|
||||
def test_parse_job_dates_from_epoch_ms():
|
||||
item = parse_job(_raw_job())
|
||||
assert item["datePublished"] == "2026-03-23T05:00:00.000Z"
|
||||
assert item["age"] == "3 days ago"
|
||||
|
||||
|
||||
def test_parse_job_respects_base_url():
|
||||
item = parse_job(_raw_job(), base_url="https://uk.indeed.com")
|
||||
assert item["jobUrl"] == "https://uk.indeed.com/viewjob?jk=abc123"
|
||||
assert item["companyUrl"] == "https://uk.indeed.com/cmp/Acme-Corp"
|
||||
|
||||
|
||||
def test_extract_blob_anchors_on_assignment_not_first_occurrence():
|
||||
# Decoy mention precedes the real assignment; the extractor must skip it.
|
||||
html = (
|
||||
'<script>var providers=["mosaic-provider-jobcards"];</script>'
|
||||
'<script>window.mosaic.providerData["mosaic-provider-jobcards"]='
|
||||
'{"metaData":{"mosaicProviderJobCardsModel":{"results":'
|
||||
'[{"jobkey":"k1"},{"jobkey":"k2"}]}}};</script>'
|
||||
)
|
||||
blob = extract_jobcards_blob(html)
|
||||
results = job_results(blob)
|
||||
assert [r["jobkey"] for r in results] == ["k1", "k2"]
|
||||
|
||||
|
||||
def test_extract_blob_missing_returns_none():
|
||||
assert extract_jobcards_blob("<html>just a moment...</html>") is None
|
||||
assert job_results(None) == []
|
||||
|
||||
|
||||
# --- fixture-pinned (real captured blob) -----------------------------------
|
||||
|
||||
|
||||
def test_fixture_blob_parses_into_items():
|
||||
fixture = _FIXTURE_DIR / "sample_jobcards.json"
|
||||
blob = json.loads(fixture.read_text())
|
||||
results = job_results(blob)
|
||||
assert len(results) == 3
|
||||
for raw in results:
|
||||
item = parse_job(raw)
|
||||
assert isinstance(item["jobKey"], str) and item["jobKey"]
|
||||
assert item["title"]
|
||||
assert item["jobUrl"].startswith("https://www.indeed.com/viewjob?jk=")
|
||||
assert "salaryText" in item["salary"]
|
||||
|
||||
|
||||
# --- detail page (parse_job_detail) ----------------------------------------
|
||||
|
||||
|
||||
def test_parse_job_detail_extracts_description_and_fields():
|
||||
html = (_FIXTURE_DIR / "sample_viewjob.html").read_text()
|
||||
detail = parse_job_detail(html)
|
||||
assert detail["descriptionHtml"].startswith("<div>")
|
||||
assert "Data Analyst" in detail["descriptionText"]
|
||||
assert "<div>" not in detail["descriptionText"] # tags stripped
|
||||
assert detail["title"]
|
||||
assert detail["company"]
|
||||
assert detail["formattedLocation"]
|
||||
assert detail["jobTypes"] == ["Full-time"]
|
||||
assert detail["benefits"] == ["401(k)", "Health insurance"]
|
||||
assert detail["isRemote"] is True
|
||||
assert detail["remoteType"] == "HYBRID"
|
||||
sal = detail["salary"]
|
||||
assert (sal["salaryMin"], sal["salaryMax"], sal["period"]) == (60000, 90000, "year")
|
||||
|
||||
|
||||
def test_parse_job_detail_reads_rootprops_shape():
|
||||
# Live /viewjob assigns window._rootProps (JSON) with the model under
|
||||
# preloadedVJData; window._initialData is now a non-JSON JS literal that
|
||||
# references other globals and must be skipped, not misparsed.
|
||||
html = (
|
||||
"<html><script>window._initialData = { ssr: true, "
|
||||
"viewJobClientSideModel: window._rootProps.preloadedVJData };</script>"
|
||||
'<script>window._rootProps = {"url":"x","preloadedVJData":'
|
||||
'{"jobInfoWrapperModel":{"jobInfoModel":'
|
||||
'{"sanitizedJobDescription":"<p>Full JD text</p>",'
|
||||
'"jobInfoHeaderModel":{"jobTitle":"Data Analyst",'
|
||||
'"companyName":"Acme"}}}}};</script></html>'
|
||||
)
|
||||
detail = parse_job_detail(html)
|
||||
assert detail["title"] == "Data Analyst"
|
||||
assert detail["company"] == "Acme"
|
||||
assert detail["descriptionHtml"] == "<p>Full JD text</p>"
|
||||
assert detail["descriptionText"] == "Full JD text"
|
||||
|
||||
|
||||
def test_parse_job_detail_omits_blank_fields():
|
||||
# A page with a header but no salary/description must not emit those keys,
|
||||
# so a merge won't clobber listing values with blanks.
|
||||
html = (
|
||||
"<html><script>window._initialData = "
|
||||
'{"jobInfoWrapperModel":{"jobInfoModel":{"jobInfoHeaderModel":'
|
||||
'{"jobTitle":"Analyst"}}}};</script></html>'
|
||||
)
|
||||
detail = parse_job_detail(html)
|
||||
assert detail == {"title": "Analyst"}
|
||||
|
||||
|
||||
def test_parse_job_detail_not_a_job_page_returns_empty():
|
||||
assert parse_job_detail("<html>just a moment...</html>") == {}
|
||||
|
|
@ -0,0 +1,155 @@
|
|||
"""Offline orchestration tests: pagination, dedupe, and caps via a fake session."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from urllib.parse import parse_qs, urlparse
|
||||
|
||||
import pytest
|
||||
|
||||
from app.proprietary.platforms.indeed_jobs.fetch import IndeedAccessBlockedError
|
||||
from app.proprietary.platforms.indeed_jobs.schemas import IndeedScrapeInput
|
||||
from app.proprietary.platforms.indeed_jobs.scraper import iter_indeed, scrape_indeed
|
||||
|
||||
|
||||
def _page_html(job_keys: list[str]) -> str:
|
||||
"""Wrap job keys in the ``mosaic-provider-jobcards`` assignment shape."""
|
||||
results = [
|
||||
{"jobkey": k, "displayTitle": f"Job {k}", "company": "Acme"} for k in job_keys
|
||||
]
|
||||
model = {"metaData": {"mosaicProviderJobCardsModel": {"results": results}}}
|
||||
return (
|
||||
f'window.mosaic.providerData["mosaic-provider-jobcards"]={json.dumps(model)};'
|
||||
)
|
||||
|
||||
|
||||
def _detail_html(job_key: str) -> str:
|
||||
"""A minimal /viewjob page carrying a full description for ``job_key``."""
|
||||
data = {
|
||||
"jobInfoWrapperModel": {
|
||||
"jobInfoModel": {
|
||||
"sanitizedJobDescription": f"<div>Full description for {job_key}</div>",
|
||||
"jobInfoHeaderModel": {"jobTitle": f"Detailed {job_key}"},
|
||||
}
|
||||
}
|
||||
}
|
||||
return f"<html><script>window._initialData = {json.dumps(data)};</script></html>"
|
||||
|
||||
|
||||
class _FakeSession:
|
||||
"""Returns per-``start`` search pages (or a /viewjob detail) and records URLs."""
|
||||
|
||||
def __init__(
|
||||
self, pages: dict[int, list[str]], blocked_starts: set[int] | None = None
|
||||
) -> None:
|
||||
self._pages = pages
|
||||
self._blocked = blocked_starts or set()
|
||||
self.fetched: list[str] = []
|
||||
|
||||
async def fetch_html(self, url: str, *, max_rotations: int | None = None) -> str:
|
||||
self.fetched.append(url)
|
||||
query = parse_qs(urlparse(url).query)
|
||||
if "/viewjob" in url:
|
||||
return _detail_html(query.get("jk", [""])[0])
|
||||
start = int(query.get("start", ["0"])[0])
|
||||
if start in self._blocked:
|
||||
raise IndeedAccessBlockedError(f"gated at start={start}")
|
||||
return _page_html(self._pages.get(start, []))
|
||||
|
||||
|
||||
async def _collect(input_model, session) -> list[dict]:
|
||||
return [item async for item in iter_indeed(input_model, session)]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_dedupes_within_page():
|
||||
session = _FakeSession({0: ["k1", "k2", "k2", "k3"]})
|
||||
items = await _collect(
|
||||
IndeedScrapeInput(queries=["dev"], maxItemsPerQuery=100), session
|
||||
)
|
||||
assert [i["jobKey"] for i in items] == ["k1", "k2", "k3"]
|
||||
assert all(i["scrapedAt"] for i in items) # stamped by the orchestrator
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_does_not_fetch_deeper_pages():
|
||||
# First page only; ``start>=10`` must never be requested.
|
||||
session = _FakeSession({0: ["k1", "k2"], 10: ["k3"]})
|
||||
items = await _collect(
|
||||
IndeedScrapeInput(queries=["dev"], maxItemsPerQuery=100), session
|
||||
)
|
||||
assert [i["jobKey"] for i in items] == ["k1", "k2"]
|
||||
assert all("start=" not in u for u in session.fetched)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_page_block_propagates():
|
||||
# Nothing yielded before the block, so it surfaces as an error.
|
||||
session = _FakeSession({}, blocked_starts={0})
|
||||
with pytest.raises(IndeedAccessBlockedError):
|
||||
await _collect(IndeedScrapeInput(queries=["dev"]), session)
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_respects_max_items_per_query():
|
||||
session = _FakeSession({0: ["k1", "k2", "k3", "k4"]})
|
||||
items = await _collect(
|
||||
IndeedScrapeInput(queries=["dev"], maxItemsPerQuery=2), session
|
||||
)
|
||||
assert [i["jobKey"] for i in items] == ["k1", "k2"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_global_dedupe_across_queries():
|
||||
# Both queries hit page 0 (same fake pages) and return the same keys.
|
||||
session = _FakeSession({0: ["k1", "k2"]})
|
||||
items = await _collect(
|
||||
IndeedScrapeInput(queries=["dev", "engineer"], maxItemsPerQuery=100), session
|
||||
)
|
||||
assert [i["jobKey"] for i in items] == ["k1", "k2"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_start_urls_scrape_search_and_job_url_detail():
|
||||
session = _FakeSession({0: ["k1"]})
|
||||
input_model = IndeedScrapeInput(
|
||||
startUrls=[
|
||||
{"url": "https://www.indeed.com/jobs?q=dev"},
|
||||
{"url": "https://www.indeed.com/viewjob?jk=abc"},
|
||||
],
|
||||
maxItemsPerQuery=100,
|
||||
)
|
||||
items = await _collect(input_model, session)
|
||||
assert len(items) == 2
|
||||
search_item, job_item = items
|
||||
assert search_item["jobKey"] == "k1"
|
||||
# The /viewjob URL is scraped from its detail page alone.
|
||||
assert job_item["jobUrl"].endswith("jk=abc")
|
||||
assert job_item["title"] == "Detailed abc"
|
||||
assert "Full description for abc" in job_item["descriptionText"]
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_scrape_job_details_enriches_listing_items():
|
||||
session = _FakeSession({0: ["k1", "k2"]})
|
||||
items = await _collect(
|
||||
IndeedScrapeInput(queries=["dev"], maxItemsPerQuery=100, scrapeJobDetails=True),
|
||||
session,
|
||||
)
|
||||
assert [i["jobKey"] for i in items] == ["k1", "k2"]
|
||||
for it in items:
|
||||
assert it["descriptionHtml"].startswith("<div>Full description for")
|
||||
assert "Full description for" in it["descriptionText"]
|
||||
# One extra /viewjob load per listing item.
|
||||
assert sum("/viewjob" in u for u in session.fetched) == 2
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_scrape_indeed_limit_with_injected_session():
|
||||
session = _FakeSession({0: ["k1", "k2", "k3"]})
|
||||
items = await scrape_indeed(
|
||||
IndeedScrapeInput(queries=["dev"], maxItemsPerQuery=100),
|
||||
limit=2,
|
||||
session=session,
|
||||
)
|
||||
assert [i["jobKey"] for i in items] == ["k1", "k2"]
|
||||
|
|
@ -0,0 +1,87 @@
|
|||
"""Offline tests for Indeed URL classification and search-URL building."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from urllib.parse import parse_qs, urlparse
|
||||
|
||||
from app.proprietary.platforms.indeed_jobs.url_resolver import (
|
||||
build_search_url,
|
||||
country_domain,
|
||||
resolve_url,
|
||||
)
|
||||
|
||||
|
||||
def test_resolve_search_url():
|
||||
r = resolve_url(
|
||||
"https://www.indeed.com/jobs?q=software+engineer&l=Remote&sort=date"
|
||||
)
|
||||
assert r is not None
|
||||
assert r.kind == "search"
|
||||
assert r.value == "software engineer"
|
||||
assert r.location == "Remote"
|
||||
assert r.domain == "www.indeed.com"
|
||||
assert r.params.get("sort") == "date"
|
||||
|
||||
|
||||
def test_resolve_company_url():
|
||||
r = resolve_url("https://www.indeed.com/cmp/Google/jobs")
|
||||
assert r is not None
|
||||
assert r.kind == "company"
|
||||
assert r.value == "Google"
|
||||
|
||||
|
||||
def test_resolve_viewjob_url():
|
||||
r = resolve_url("https://uk.indeed.com/viewjob?jk=abc123&from=serp")
|
||||
assert r is not None
|
||||
assert r.kind == "job"
|
||||
assert r.value == "abc123"
|
||||
assert r.domain == "uk.indeed.com"
|
||||
|
||||
|
||||
def test_resolve_country_subdomain_host():
|
||||
r = resolve_url("https://de.indeed.com/jobs?q=entwickler")
|
||||
assert r is not None
|
||||
assert r.kind == "search"
|
||||
assert r.domain == "de.indeed.com"
|
||||
|
||||
|
||||
def test_resolve_rejects_non_indeed():
|
||||
assert resolve_url("https://www.linkedin.com/jobs?q=dev") is None
|
||||
assert resolve_url("https://notindeed.com.evil.com/jobs") is None
|
||||
|
||||
|
||||
def test_country_domain_map():
|
||||
assert country_domain("us") == "www.indeed.com"
|
||||
assert country_domain("gb") == "uk.indeed.com"
|
||||
assert country_domain("de") == "de.indeed.com"
|
||||
assert country_domain("") == "www.indeed.com"
|
||||
|
||||
|
||||
def test_build_search_url_basic():
|
||||
url = build_search_url(
|
||||
"data analyst",
|
||||
country="us",
|
||||
location="New York, NY",
|
||||
sort="date",
|
||||
start=20,
|
||||
)
|
||||
parsed = urlparse(url)
|
||||
qs = parse_qs(parsed.query)
|
||||
assert parsed.netloc == "www.indeed.com"
|
||||
assert parsed.path == "/jobs"
|
||||
assert qs["q"] == ["data analyst"]
|
||||
assert qs["l"] == ["New York, NY"]
|
||||
assert qs["sort"] == ["date"]
|
||||
assert qs["start"] == ["20"]
|
||||
|
||||
|
||||
def test_build_search_url_remote_keyword_fallback_and_jobtype():
|
||||
url = build_search_url(
|
||||
"developer", country="gb", remote="remote", job_type="fulltime", from_days=7
|
||||
)
|
||||
parsed = urlparse(url)
|
||||
qs = parse_qs(parsed.query)
|
||||
assert parsed.netloc == "uk.indeed.com"
|
||||
assert qs["q"] == ["developer remote"]
|
||||
assert qs["jt"] == ["fulltime"]
|
||||
assert qs["fromage"] == ["7"]
|
||||
|
|
@ -1,7 +1,7 @@
|
|||
"""Scraper tools: one MCP surface per SurfSense platform capability.
|
||||
|
||||
Web crawl, Google Search, Reddit, YouTube, and Google Maps each get a tool that
|
||||
maps a natural-language request to the workspace's scraper. Two run-history tools
|
||||
Web crawl, Google Search, Reddit, YouTube, Google Maps, Indeed, and Amazon each
|
||||
get a tool that maps a natural-language request to the workspace's scraper. Two run-history tools
|
||||
list and fetch past runs, so a large result truncated inline can be retrieved in
|
||||
full later. Each platform lives in its own module under platforms/.
|
||||
"""
|
||||
|
|
@ -17,6 +17,7 @@ from .platforms import (
|
|||
amazon,
|
||||
google_maps,
|
||||
google_search,
|
||||
indeed,
|
||||
instagram,
|
||||
reddit,
|
||||
tiktok,
|
||||
|
|
@ -32,6 +33,7 @@ _REGISTRARS = (
|
|||
instagram,
|
||||
tiktok,
|
||||
google_maps,
|
||||
indeed,
|
||||
amazon,
|
||||
run_history,
|
||||
)
|
||||
|
|
|
|||
135
surfsense_mcp/mcp_server/features/scrapers/platforms/indeed.py
Normal file
135
surfsense_mcp/mcp_server/features/scrapers/platforms/indeed.py
Normal file
|
|
@ -0,0 +1,135 @@
|
|||
"""Indeed scraper tool."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from typing import Annotated, Literal
|
||||
|
||||
from mcp.server.fastmcp import FastMCP
|
||||
from pydantic import Field
|
||||
|
||||
from ....core.client import SurfSenseClient
|
||||
from ....core.rendering import ResponseFormatParam
|
||||
from ....core.workspace_context import WorkspaceContext, WorkspaceParam
|
||||
from ..annotations import SCRAPE
|
||||
from ..capability import run_scraper
|
||||
|
||||
IndeedSort = Literal["relevance", "date"]
|
||||
IndeedJobType = Literal[
|
||||
"fulltime",
|
||||
"parttime",
|
||||
"contract",
|
||||
"internship",
|
||||
"temporary",
|
||||
"permanent",
|
||||
"seasonal",
|
||||
"freelance",
|
||||
]
|
||||
IndeedLevel = Literal["entry_level", "mid_level", "senior_level"]
|
||||
IndeedRemote = Literal["remote", "hybrid"]
|
||||
|
||||
|
||||
def register(mcp: FastMCP, client: SurfSenseClient, context: WorkspaceContext) -> None:
|
||||
"""Register the Indeed tool."""
|
||||
|
||||
@mcp.tool(
|
||||
name="surfsense_indeed_scrape",
|
||||
title="Search or scrape Indeed jobs",
|
||||
annotations=SCRAPE,
|
||||
structured_output=False,
|
||||
)
|
||||
async def indeed_scrape(
|
||||
urls: Annotated[
|
||||
list[str] | None,
|
||||
Field(
|
||||
description="Indeed URLs: a search page "
|
||||
"('https://www.indeed.com/jobs?q=data+analyst'), a company jobs "
|
||||
"page ('/cmp/<slug>/jobs'), or a single job ('/viewjob?jk=...'). "
|
||||
"Provide urls OR search_queries."
|
||||
),
|
||||
] = None,
|
||||
search_queries: Annotated[
|
||||
list[str] | None,
|
||||
Field(
|
||||
description="Job search terms, e.g. ['data analyst', 'ml engineer']. "
|
||||
"Provide search_queries OR urls."
|
||||
),
|
||||
] = None,
|
||||
country: Annotated[
|
||||
str,
|
||||
Field(description="Country code selecting the Indeed domain, e.g. 'us', 'gb'."),
|
||||
] = "us",
|
||||
location: Annotated[
|
||||
str | None,
|
||||
Field(description="Where to search, e.g. 'Remote', 'New York, NY'."),
|
||||
] = None,
|
||||
radius: Annotated[
|
||||
int | None,
|
||||
Field(description="Search radius in miles/km around location."),
|
||||
] = None,
|
||||
job_type: Annotated[
|
||||
IndeedJobType | None,
|
||||
Field(description="Employment type filter."),
|
||||
] = None,
|
||||
level: Annotated[
|
||||
IndeedLevel | None,
|
||||
Field(description="Experience level filter."),
|
||||
] = None,
|
||||
remote: Annotated[
|
||||
IndeedRemote | None,
|
||||
Field(description="Work model filter: remote or hybrid."),
|
||||
] = None,
|
||||
from_days: Annotated[
|
||||
int | None,
|
||||
Field(description="Only return jobs posted within the last N days."),
|
||||
] = None,
|
||||
sort: Annotated[
|
||||
IndeedSort, Field(description="Result ordering: relevance or date.")
|
||||
] = "relevance",
|
||||
scrape_job_details: Annotated[
|
||||
bool,
|
||||
Field(
|
||||
description="True fetches each job's detail page for the full "
|
||||
"description (slower); False returns the listing snippet only."
|
||||
),
|
||||
] = False,
|
||||
max_items: Annotated[
|
||||
int, Field(ge=1, description="Maximum jobs to return in total.")
|
||||
] = 25,
|
||||
max_items_per_query: Annotated[
|
||||
int, Field(ge=0, description="Max jobs per search/company target.")
|
||||
] = 25,
|
||||
workspace: WorkspaceParam = None,
|
||||
response_format: ResponseFormatParam = "markdown",
|
||||
) -> str:
|
||||
"""Search or scrape public Indeed job postings.
|
||||
|
||||
Use this for ANY Indeed job research — openings for a role, who is hiring
|
||||
at a company, salaries for a title in a location, or remote roles —
|
||||
instead of a generic web search. Returns jobs with title, company,
|
||||
location, salary, job types, and description; set scrape_job_details for
|
||||
the full description per job.
|
||||
Example: search_queries=['data analyst'], location='Remote', max_items=30.
|
||||
"""
|
||||
return await run_scraper(
|
||||
client,
|
||||
context,
|
||||
platform="indeed",
|
||||
verb="scrape",
|
||||
payload={
|
||||
"urls": urls,
|
||||
"search_queries": search_queries,
|
||||
"country": country,
|
||||
"location": location,
|
||||
"radius": radius,
|
||||
"job_type": job_type,
|
||||
"level": level,
|
||||
"remote": remote,
|
||||
"from_days": from_days,
|
||||
"sort": sort,
|
||||
"scrape_job_details": scrape_job_details,
|
||||
"max_items": max_items,
|
||||
"max_items_per_query": max_items_per_query,
|
||||
},
|
||||
workspace=workspace,
|
||||
response_format=response_format,
|
||||
)
|
||||
|
|
@ -32,6 +32,7 @@ EXPECTED_TOOLS = {
|
|||
"surfsense_amazon_scrape",
|
||||
"surfsense_instagram_scrape",
|
||||
"surfsense_instagram_details",
|
||||
"surfsense_indeed_scrape",
|
||||
"surfsense_list_scraper_runs",
|
||||
"surfsense_get_scraper_run",
|
||||
# knowledge-base management
|
||||
|
|
|
|||
|
|
@ -38,7 +38,8 @@ def build_server(settings: Settings) -> tuple[FastMCP, SurfSenseClient]:
|
|||
"task involves Reddit (posts, comments, finding subreddits or "
|
||||
"communities), YouTube (videos, transcripts, comments), Instagram "
|
||||
"(posts, reels, profile details), TikTok (videos by hashtag, "
|
||||
"search, or URL), Google Maps (places, reviews), Google Search "
|
||||
"search, or URL), Google Maps (places, reviews), Indeed (job "
|
||||
"postings by role, company, or location), Google Search "
|
||||
"results, or reading "
|
||||
"specific web pages. Scraper results are persisted as runs; if an "
|
||||
"inline result is truncated, fetch it in full with "
|
||||
|
|
|
|||
|
|
@ -12,7 +12,7 @@ Connectors bring data into SurfSense — either as searchable knowledge or as li
|
|||
<Card
|
||||
icon={<Zap />}
|
||||
title="Native Connectors"
|
||||
description="SurfSense's built-in scraper APIs: Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, and Web Crawl — usable in chat, the API Playground, or via REST"
|
||||
description="SurfSense's built-in scraper APIs: Reddit, YouTube, Instagram, TikTok, Google Maps, Google Search, Indeed, Amazon, and Web Crawl — usable in chat, the API Playground, or via REST"
|
||||
href="/docs/connectors/native"
|
||||
/>
|
||||
<Card
|
||||
|
|
|
|||
51
surfsense_web/content/docs/connectors/native/indeed.mdx
Normal file
51
surfsense_web/content/docs/connectors/native/indeed.mdx
Normal file
|
|
@ -0,0 +1,51 @@
|
|||
---
|
||||
title: Indeed
|
||||
description: Scrape public Indeed job postings by search, company, or URL
|
||||
---
|
||||
|
||||
The Indeed scraper pulls structured public job postings from Indeed. Give it search terms and/or URLs (a search page, a company jobs page, or a single job), and it returns jobs with title, company, location, salary, job types, benefits, remote/hybrid flag, posting age, and description.
|
||||
|
||||
## Endpoint
|
||||
|
||||
```bash
|
||||
POST /api/v1/workspaces/{workspace_id}/scrapers/indeed/scrape
|
||||
```
|
||||
|
||||
## Inputs
|
||||
|
||||
At least one of `urls` or `search_queries` is required.
|
||||
|
||||
| Field | Default | Description |
|
||||
|-------|---------|-------------|
|
||||
| `urls` | — | Indeed URLs: a search page (`/jobs?q=&l=`), a company jobs page (`/cmp/<slug>/jobs`), or a single job (`/viewjob?jk=`) (max 20 sources per call, combined with queries) |
|
||||
| `search_queries` | — | Job search terms, e.g. `["data analyst"]` |
|
||||
| `country` | `us` | Country code selecting the Indeed domain, e.g. `us`, `gb`, `de` |
|
||||
| `location` | — | Where to search, e.g. `Remote`, `New York, NY` |
|
||||
| `radius` | — | Search radius in miles/km around `location` |
|
||||
| `job_type` | — | `fulltime`, `parttime`, `contract`, `internship`, `temporary`, `permanent`, `seasonal`, or `freelance` |
|
||||
| `level` | — | `entry_level`, `mid_level`, or `senior_level` |
|
||||
| `remote` | — | `remote` or `hybrid` |
|
||||
| `from_days` | — | Only return jobs posted within the last N days |
|
||||
| `sort` | `relevance` | `relevance` or `date` |
|
||||
| `scrape_job_details` | `false` | Fetch each job's detail page for the full description (slower: one extra page load per job) |
|
||||
| `max_items` | `25` | Max total jobs returned across all sources (hard cap 100) |
|
||||
| `max_items_per_query` | `25` | Max jobs per search/company target |
|
||||
|
||||
## Example
|
||||
|
||||
```bash
|
||||
curl -X POST "$BASE_URL/api/v1/workspaces/1/scrapers/indeed/scrape" \
|
||||
-H "Authorization: Bearer $SURFSENSE_API_KEY" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"search_queries": ["data analyst"],
|
||||
"location": "Remote",
|
||||
"remote": "remote",
|
||||
"sort": "date",
|
||||
"max_items": 30
|
||||
}'
|
||||
```
|
||||
|
||||
The response is `{ "items": [...] }` — one item per job posting. Billing is per returned job.
|
||||
|
||||
For the full input and output JSON schemas and generated code snippets in your language, open **API Playground → Indeed → Scrape** in your workspace.
|
||||
|
|
@ -1,6 +1,6 @@
|
|||
---
|
||||
title: Native Connectors
|
||||
description: SurfSense's built-in scraper APIs for Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, and the web
|
||||
description: SurfSense's built-in scraper APIs for Reddit, YouTube, Instagram, TikTok, Google Maps, Google Search, Indeed, Amazon, and the web
|
||||
---
|
||||
|
||||
import { Card, Cards } from 'fumadocs-ui/components/card';
|
||||
|
|
@ -38,6 +38,11 @@ Native connectors are SurfSense's own scraper APIs — built into the platform,
|
|||
description="Structured SERPs: organic results, people-also-ask, AI overviews"
|
||||
href="/docs/connectors/native/google-search"
|
||||
/>
|
||||
<Card
|
||||
title="Indeed"
|
||||
description="Job postings by search, company, or URL — with salary and full description"
|
||||
href="/docs/connectors/native/indeed"
|
||||
/>
|
||||
<Card
|
||||
title="Amazon"
|
||||
description="Public product data: prices, ratings, offers, sellers, and best-seller ranks"
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@
|
|||
"tiktok",
|
||||
"google-maps",
|
||||
"google-search",
|
||||
"indeed",
|
||||
"amazon",
|
||||
"web-crawl"
|
||||
],
|
||||
|
|
|
|||
|
|
@ -7,7 +7,7 @@ import { Tab, Tabs } from 'fumadocs-ui/components/tabs';
|
|||
|
||||
# SurfSense MCP Server
|
||||
|
||||
The SurfSense MCP server exposes your workspace to any [Model Context Protocol](https://modelcontextprotocol.io/) client. Your agent gets 25 native, typed tools: every scraper (Reddit, YouTube, Instagram, TikTok, Amazon, Google Maps, Google Search, web crawl), full knowledge-base access (search, read, add, upload, update, delete), and a workspace selector.
|
||||
The SurfSense MCP server exposes your workspace to any [Model Context Protocol](https://modelcontextprotocol.io/) client. Your agent gets 26 native, typed tools: every scraper (Reddit, YouTube, Instagram, TikTok, Google Maps, Google Search, Indeed, Amazon, web crawl), full knowledge-base access (search, read, add, upload, update, delete), and a workspace selector.
|
||||
|
||||
Connect it two ways: the **hosted** server at `https://mcp.surfsense.com/mcp` (nothing to install — just an API key), or run it yourself over **stdio** against any SurfSense backend, cloud or self-hosted.
|
||||
|
||||
|
|
@ -264,7 +264,7 @@ For self-host (stdio), all settings are environment variables passed by the clie
|
|||
| Group | Tools |
|
||||
|-------|-------|
|
||||
| Workspaces | `surfsense_list_workspaces`, `surfsense_select_workspace` |
|
||||
| Scrapers | `surfsense_reddit_scrape`, `surfsense_youtube_scrape`, `surfsense_youtube_comments`, `surfsense_instagram_scrape`, `surfsense_instagram_details`, `surfsense_tiktok_scrape`, `surfsense_tiktok_comments`, `surfsense_tiktok_user_search`, `surfsense_tiktok_trending`, `surfsense_google_maps_scrape`, `surfsense_google_maps_reviews`, `surfsense_google_search`, `surfsense_web_crawl`, `surfsense_list_scraper_runs`, `surfsense_get_scraper_run` |
|
||||
| Scrapers | `surfsense_reddit_scrape`, `surfsense_youtube_scrape`, `surfsense_youtube_comments`, `surfsense_instagram_scrape`, `surfsense_instagram_details`, `surfsense_tiktok_scrape`, `surfsense_tiktok_comments`, `surfsense_tiktok_user_search`, `surfsense_tiktok_trending`, `surfsense_google_maps_scrape`, `surfsense_google_maps_reviews`, `surfsense_google_search`, `surfsense_indeed_scrape`, `surfsense_web_crawl`, `surfsense_list_scraper_runs`, `surfsense_get_scraper_run` |
|
||||
| Knowledge base | `surfsense_search_knowledge_base`, `surfsense_list_documents`, `surfsense_get_document`, `surfsense_add_document`, `surfsense_upload_file`, `surfsense_update_document`, `surfsense_delete_document` |
|
||||
|
||||
Usage is billed exactly like the REST API — scraper tools are metered per returned item, and every call is recorded under **API Playground → Runs**.
|
||||
|
|
|
|||
|
|
@ -44,6 +44,7 @@ const PUBLIC_ROUTE_PREFIXES = [
|
|||
"/youtube",
|
||||
"/google-maps",
|
||||
"/google-search",
|
||||
"/indeed",
|
||||
"/web-crawl",
|
||||
"/amazon",
|
||||
];
|
||||
|
|
|
|||
|
|
@ -306,6 +306,7 @@ export const googleMaps: ConnectorPageContent = {
|
|||
{ label: "Instagram API", href: "/instagram" },
|
||||
{ label: "SERP API", href: "/google-search" },
|
||||
{ label: "Web Crawl API", href: "/web-crawl" },
|
||||
{ label: "Indeed API", href: "/indeed" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
{ label: "Read the docs", href: "/docs" },
|
||||
],
|
||||
|
|
|
|||
|
|
@ -260,6 +260,7 @@ export const googleSearch: ConnectorPageContent = {
|
|||
{ label: "Google Maps API", href: "/google-maps" },
|
||||
{ label: "Reddit API", href: "/reddit" },
|
||||
{ label: "Instagram API", href: "/instagram" },
|
||||
{ label: "Indeed API", href: "/indeed" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
{ label: "Read the docs", href: "/docs" },
|
||||
],
|
||||
|
|
|
|||
327
surfsense_web/lib/connectors-marketing/indeed.tsx
Normal file
327
surfsense_web/lib/connectors-marketing/indeed.tsx
Normal file
|
|
@ -0,0 +1,327 @@
|
|||
import { IconBriefcase } from "@tabler/icons-react";
|
||||
import type { ConnectorPageContent } from "./types";
|
||||
|
||||
export const indeed: ConnectorPageContent = {
|
||||
slug: "indeed",
|
||||
name: "Indeed",
|
||||
icon: IconBriefcase,
|
||||
|
||||
metaTitle: "Indeed Scraper API for Jobs and Hiring Data | SurfSense",
|
||||
metaDescription:
|
||||
"Scrape public Indeed job postings with the SurfSense Indeed Scraper API: titles, companies, salaries, and full descriptions by search or company. No Indeed API. Start free.",
|
||||
keywords: [
|
||||
"indeed scraper",
|
||||
"indeed scraper api",
|
||||
"indeed api",
|
||||
"indeed jobs api",
|
||||
"scrape indeed",
|
||||
"indeed job scraper",
|
||||
"job posting scraper",
|
||||
"salary data api",
|
||||
"hiring data api",
|
||||
"indeed mcp",
|
||||
"labor market data",
|
||||
"recruiting data tool",
|
||||
],
|
||||
|
||||
h1: "Indeed Scraper API for Job Postings and Hiring Data",
|
||||
heroLede:
|
||||
"The SurfSense Indeed API extracts public job postings, salaries, companies, and full descriptions by search query, company page, or job URL, without Indeed's official API. Give your AI agents a live feed of who is hiring, for what, at what pay, so you track the labor market as it moves.",
|
||||
|
||||
transcript: {
|
||||
prompt: "Find remote data analyst roles posted this week and what they pay",
|
||||
toolCall:
|
||||
'indeed.scrape({ search_queries: ["data analyst"], location: "Remote",\n remote: "remote", from_days: 7, sort: "date", max_items: 30 })',
|
||||
rows: [
|
||||
{
|
||||
primary: "Senior Data Analyst · Acme Corp",
|
||||
secondary: "Remote (US) · $120k–$145k/year · posted 2 days ago",
|
||||
tag: "salary listed",
|
||||
},
|
||||
{
|
||||
primary: "Data Analyst, Growth · Globex",
|
||||
secondary: "Remote · $95k–$110k/year · Indeed Apply",
|
||||
tag: "buying signal",
|
||||
},
|
||||
{
|
||||
primary: "Marketing Data Analyst · Initech",
|
||||
secondary: "Remote (US) · estimated $88k–$102k · posted today",
|
||||
tag: "new today",
|
||||
},
|
||||
],
|
||||
resultSummary: "30 jobs · 22 with salary · surfaced in 3.4s",
|
||||
},
|
||||
|
||||
extractIntro:
|
||||
"Every call returns structured job items. Point the API at a search query, an Indeed search or company page, or a single job URL, and set scrape_job_details for the full description per job.",
|
||||
extractFields: [
|
||||
{
|
||||
label: "Job",
|
||||
description: "Title, job key, listing URL, apply URL, and whether Indeed Apply is enabled.",
|
||||
},
|
||||
{
|
||||
label: "Company",
|
||||
description: "Company name, profile URL, star rating, and review count where available.",
|
||||
},
|
||||
{
|
||||
label: "Location",
|
||||
description: "Formatted location, city, state, postal code, country, and remote or hybrid flags.",
|
||||
},
|
||||
{
|
||||
label: "Salary",
|
||||
description:
|
||||
"Pay text, min and max bounds, currency, period, and whether the figure is an Indeed estimate.",
|
||||
},
|
||||
{
|
||||
label: "Description",
|
||||
description:
|
||||
"Listing snippet by default; the full text and HTML description with scrape_job_details.",
|
||||
},
|
||||
{
|
||||
label: "Signals",
|
||||
description:
|
||||
"Job types, benefits, sponsored, urgently hiring, new, and expired flags, plus post age.",
|
||||
},
|
||||
],
|
||||
|
||||
useCasesHeading: "What teams do with the Indeed API",
|
||||
useCases: [
|
||||
{
|
||||
title: "Competitor hiring intelligence",
|
||||
description:
|
||||
"Track what your competitors are hiring for and where. A spike in sales or ML roles is a roadmap signal months before it ships. Feed the stream to an agent that flags the moves that matter.",
|
||||
},
|
||||
{
|
||||
title: "Salary and compensation benchmarking",
|
||||
description:
|
||||
"Pull real posted salaries for a title in a location and benchmark your own bands against the live market, instead of a survey that is a year stale.",
|
||||
},
|
||||
{
|
||||
title: "Labor market and sector research",
|
||||
description:
|
||||
"Measure hiring demand for a role, skill, or sector over time. Turn thousands of postings into a demand index your analysts and clients can act on.",
|
||||
},
|
||||
{
|
||||
title: "Recruiting and lead sourcing",
|
||||
description:
|
||||
"Find companies actively hiring for a role and reach them while the need is hot. Job postings are a public, timely buying signal for staffing and B2B sales.",
|
||||
},
|
||||
],
|
||||
|
||||
comparison: {
|
||||
heading: "An Indeed API alternative built for agents",
|
||||
intro:
|
||||
"Indeed retired its public Publisher jobs API and gates data behind partner programs. If you cannot get access or need clean structured jobs now, here is how SurfSense compares.",
|
||||
columnLabel: "DIY Indeed scraping",
|
||||
rows: [
|
||||
{
|
||||
feature: "Access",
|
||||
official: "Publisher API retired; partner-gated and approval-only",
|
||||
surfsense: "One API key; scrape public postings without an approval process",
|
||||
},
|
||||
{
|
||||
feature: "Anti-bot",
|
||||
official: "You fight Cloudflare, fingerprinting, and CAPTCHAs yourself",
|
||||
surfsense: "Warmed, rotated sessions managed for you; no proxy plumbing",
|
||||
},
|
||||
{
|
||||
feature: "Pricing",
|
||||
official: "Proxy, browser, and maintenance costs you own",
|
||||
surfsense: "Pay per job returned, with a free tier to start",
|
||||
},
|
||||
{
|
||||
feature: "Descriptions",
|
||||
official: "Extra page fetch and parsing you build and maintain",
|
||||
surfsense: "Full description per job with one scrape_job_details flag",
|
||||
},
|
||||
{
|
||||
feature: "Agent-ready",
|
||||
official: "No; you build the harness yourself",
|
||||
surfsense: "MCP server exposes indeed.scrape as a native tool",
|
||||
},
|
||||
],
|
||||
},
|
||||
|
||||
api: {
|
||||
platform: "indeed",
|
||||
verb: "scrape",
|
||||
mcpTool: "indeed.scrape",
|
||||
requestBody: {
|
||||
search_queries: ["data analyst"],
|
||||
location: "Remote",
|
||||
remote: "remote",
|
||||
from_days: 7,
|
||||
sort: "date",
|
||||
max_items: 30,
|
||||
},
|
||||
},
|
||||
|
||||
schema: {
|
||||
requestNote:
|
||||
"Provide at least one source: urls or search_queries. Up to 20 sources per call.",
|
||||
request: [
|
||||
{
|
||||
name: "urls",
|
||||
type: "string[]",
|
||||
defaultValue: "[]",
|
||||
description:
|
||||
"Indeed URLs: a search page (/jobs?q=&l=), a company jobs page (/cmp/<slug>/jobs), or a single job (/viewjob?jk=...). Max 20.",
|
||||
},
|
||||
{
|
||||
name: "search_queries",
|
||||
type: "string[]",
|
||||
defaultValue: "[]",
|
||||
description:
|
||||
"Job search terms. Each returns up to max_items_per_query results, shaped by the filters below. Max 20.",
|
||||
},
|
||||
{
|
||||
name: "country",
|
||||
type: "string",
|
||||
defaultValue: '"us"',
|
||||
description: "Country code selecting the Indeed domain, e.g. 'us', 'gb', 'de'.",
|
||||
},
|
||||
{
|
||||
name: "location",
|
||||
type: "string",
|
||||
description: "Where to search, e.g. 'Remote', 'New York, NY'.",
|
||||
},
|
||||
{
|
||||
name: "radius",
|
||||
type: "integer",
|
||||
description: "Search radius in miles or km around location.",
|
||||
},
|
||||
{
|
||||
name: "job_type",
|
||||
type: "string",
|
||||
description: "Employment type: fulltime, parttime, contract, internship, and more.",
|
||||
},
|
||||
{
|
||||
name: "level",
|
||||
type: "string",
|
||||
description: "Experience level: entry_level, mid_level, or senior_level.",
|
||||
},
|
||||
{
|
||||
name: "remote",
|
||||
type: "string",
|
||||
description: "Work model filter: remote or hybrid.",
|
||||
},
|
||||
{
|
||||
name: "from_days",
|
||||
type: "integer",
|
||||
description: "Only return jobs posted within the last N days.",
|
||||
},
|
||||
{
|
||||
name: "sort",
|
||||
type: "string",
|
||||
defaultValue: '"relevance"',
|
||||
description: "Result ordering: relevance or date.",
|
||||
},
|
||||
{
|
||||
name: "scrape_job_details",
|
||||
type: "boolean",
|
||||
defaultValue: "false",
|
||||
description:
|
||||
"Fetch each job's detail page for the full description. Slower: one extra page load per job.",
|
||||
},
|
||||
{
|
||||
name: "max_items",
|
||||
type: "integer",
|
||||
defaultValue: "25",
|
||||
description: "Max total jobs to return across all sources. 1 to 100.",
|
||||
},
|
||||
{
|
||||
name: "max_items_per_query",
|
||||
type: "integer",
|
||||
defaultValue: "25",
|
||||
description: "Max jobs to pull per search or company target.",
|
||||
},
|
||||
],
|
||||
responseNote:
|
||||
"The response is { items: [...] } with one flat item per job. Fields Indeed omits are null. One returned job is one billable unit.",
|
||||
response: [
|
||||
{
|
||||
name: "jobKey / jobUrl / applyUrl",
|
||||
type: "string",
|
||||
description: "Indeed job key, listing URL, and third-party apply URL.",
|
||||
},
|
||||
{
|
||||
name: "title",
|
||||
type: "string",
|
||||
description: "The job title as posted.",
|
||||
},
|
||||
{
|
||||
name: "company / companyUrl",
|
||||
type: "string",
|
||||
description: "Company name and its Indeed profile URL.",
|
||||
},
|
||||
{
|
||||
name: "companyRating / companyReviewCount",
|
||||
type: "number / integer",
|
||||
description: "Employer star rating and number of reviews, where Indeed shows them.",
|
||||
},
|
||||
{
|
||||
name: "formattedLocation / isRemote / remoteType",
|
||||
type: "string / boolean",
|
||||
description: "Location string plus remote and hybrid flags.",
|
||||
},
|
||||
{
|
||||
name: "salary",
|
||||
type: "object",
|
||||
description:
|
||||
"salaryText, salaryMin, salaryMax, currency, period, and isEstimated when the pay is an Indeed estimate.",
|
||||
},
|
||||
{
|
||||
name: "jobTypes / benefits",
|
||||
type: "string[]",
|
||||
description: "Employment types and listed benefits parsed from the posting.",
|
||||
},
|
||||
{
|
||||
name: "descriptionText / descriptionHtml",
|
||||
type: "string",
|
||||
description: "Snippet by default; the full description when scrape_job_details is set.",
|
||||
},
|
||||
{
|
||||
name: "sponsored / urgentlyHiring / isNew / expired",
|
||||
type: "boolean",
|
||||
description: "Listing flags for ranking and filtering.",
|
||||
},
|
||||
{
|
||||
name: "age / datePublished / scrapedAt",
|
||||
type: "string",
|
||||
description: "Relative post age, ISO publish date, and when the job was scraped.",
|
||||
},
|
||||
],
|
||||
},
|
||||
|
||||
faq: [
|
||||
{
|
||||
question: "Is scraping Indeed legal?",
|
||||
answer:
|
||||
"SurfSense reads only public Indeed job postings, the same listings any logged-out visitor can see. It never logs in and cannot access private or applicant data. As always, review Indeed's terms and your own compliance needs before you run at scale.",
|
||||
},
|
||||
{
|
||||
question: "Does Indeed have an official jobs API?",
|
||||
answer:
|
||||
"Indeed retired its public Publisher jobs API and now gates job data behind partner and approval programs. SurfSense is an independent alternative: you call one API, or add the MCP server to your agent, and get structured public postings back.",
|
||||
},
|
||||
{
|
||||
question: "Can I get the full job description?",
|
||||
answer:
|
||||
"Yes. By default each job returns the listing snippet, which is fast. Set scrape_job_details to true and SurfSense fetches each job's detail page for the full description text and HTML, at the cost of one extra page load per job.",
|
||||
},
|
||||
{
|
||||
question: "What are the rate limits?",
|
||||
answer:
|
||||
"Each call returns up to 100 jobs across all sources, with up to 20 URLs or search queries per request. SurfSense manages the anti-bot request budget for you, so you scale reads without running proxies or a headless browser yourself.",
|
||||
},
|
||||
],
|
||||
|
||||
related: [
|
||||
{ label: "Reddit API", href: "/reddit" },
|
||||
{ label: "YouTube API", href: "/youtube" },
|
||||
{ label: "Google Maps API", href: "/google-maps" },
|
||||
{ label: "SERP API", href: "/google-search" },
|
||||
{ label: "Web Crawl API", href: "/web-crawl" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
],
|
||||
};
|
||||
|
|
@ -1,6 +1,7 @@
|
|||
import { amazon } from "./amazon";
|
||||
import { googleMaps } from "./google-maps";
|
||||
import { googleSearch } from "./google-search";
|
||||
import { indeed } from "./indeed";
|
||||
import { instagram } from "./instagram";
|
||||
import { reddit } from "./reddit";
|
||||
import { tiktok } from "./tiktok";
|
||||
|
|
@ -18,6 +19,7 @@ const CONNECTOR_LIST: ConnectorPageContent[] = [
|
|||
tiktok,
|
||||
googleMaps,
|
||||
googleSearch,
|
||||
indeed,
|
||||
amazon,
|
||||
webCrawl,
|
||||
];
|
||||
|
|
|
|||
|
|
@ -288,6 +288,7 @@ export const instagram: ConnectorPageContent = {
|
|||
{ label: "Reddit API", href: "/reddit" },
|
||||
{ label: "Google Maps API", href: "/google-maps" },
|
||||
{ label: "SERP API", href: "/google-search" },
|
||||
{ label: "Indeed API", href: "/indeed" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
],
|
||||
};
|
||||
|
|
|
|||
|
|
@ -315,6 +315,7 @@ export const reddit: ConnectorPageContent = {
|
|||
{ label: "Google Maps API", href: "/google-maps" },
|
||||
{ label: "SERP API", href: "/google-search" },
|
||||
{ label: "Web Crawl API", href: "/web-crawl" },
|
||||
{ label: "Indeed API", href: "/indeed" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
],
|
||||
};
|
||||
|
|
|
|||
|
|
@ -284,6 +284,7 @@ export const tiktok: ConnectorPageContent = {
|
|||
{ label: "Google Maps API", href: "/google-maps" },
|
||||
{ label: "SERP API", href: "/google-search" },
|
||||
{ label: "Web Crawl API", href: "/web-crawl" },
|
||||
{ label: "Indeed API", href: "/indeed" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
],
|
||||
};
|
||||
|
|
|
|||
|
|
@ -283,6 +283,7 @@ export const webCrawl: ConnectorPageContent = {
|
|||
{ label: "Google Maps API", href: "/google-maps" },
|
||||
{ label: "Reddit API", href: "/reddit" },
|
||||
{ label: "Instagram API", href: "/instagram" },
|
||||
{ label: "Indeed API", href: "/indeed" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
{ label: "Read the docs", href: "/docs" },
|
||||
],
|
||||
|
|
|
|||
|
|
@ -275,6 +275,7 @@ export const youtube: ConnectorPageContent = {
|
|||
{ label: "Google Maps API", href: "/google-maps" },
|
||||
{ label: "SERP API", href: "/google-search" },
|
||||
{ label: "Web Crawl API", href: "/web-crawl" },
|
||||
{ label: "Indeed API", href: "/indeed" },
|
||||
{ label: "SurfSense MCP Server", href: "/mcp-server" },
|
||||
],
|
||||
};
|
||||
|
|
|
|||
|
|
@ -3,6 +3,7 @@ import {
|
|||
AmazonIcon,
|
||||
GoogleMapsIcon,
|
||||
GoogleSearchIcon,
|
||||
IndeedIcon,
|
||||
InstagramIcon,
|
||||
RedditIcon,
|
||||
TikTokIcon,
|
||||
|
|
@ -90,6 +91,12 @@ export const PLAYGROUND_PLATFORMS: PlaygroundPlatform[] = [
|
|||
icon: GoogleSearchIcon,
|
||||
verbs: [{ name: "google_search.scrape", verb: "scrape", label: "Scrape" }],
|
||||
},
|
||||
{
|
||||
id: "indeed",
|
||||
label: "Indeed",
|
||||
icon: IndeedIcon,
|
||||
verbs: [{ name: "indeed.scrape", verb: "scrape", label: "Scrape" }],
|
||||
},
|
||||
{
|
||||
id: "amazon",
|
||||
label: "Amazon",
|
||||
|
|
|
|||
|
|
@ -29,4 +29,5 @@ export const InstagramIcon = brandIcon("/connectors/instagram.svg", "Instagram")
|
|||
export const TikTokIcon = brandIcon("/connectors/tiktok.svg", "TikTok");
|
||||
export const GoogleMapsIcon = brandIcon("/connectors/google-maps.svg", "Google Maps");
|
||||
export const GoogleSearchIcon = brandIcon("/connectors/google-search.svg", "Google Search");
|
||||
export const IndeedIcon = brandIcon("/connectors/indeed.svg", "Indeed");
|
||||
export const WebIcon = brandIcon("/connectors/web.svg", "Web");
|
||||
|
|
|
|||
1
surfsense_web/public/connectors/indeed.svg
Normal file
1
surfsense_web/public/connectors/indeed.svg
Normal file
|
|
@ -0,0 +1 @@
|
|||
<svg xmlns="http://www.w3.org/2000/svg" role="img" aria-label="Indeed" viewBox="0 0 24 24" width="24" height="24" fill="#003A9B"><title>Indeed</title><path d="M11.566 21.5633v-8.762c.2553.0231.5009.0346.758.0346 1.2225 0 2.3739-.3206 3.3506-.8928v9.6182c0 .8219-.1957 1.4287-.5757 1.8338-.378.4033-.8808.6049-1.491.6049-.6007 0-1.0766-.2016-1.468-.6183-.3781-.4032-.5739-1.01-.5739-1.8184zM11.589.5659c2.5447-.8929 5.4424-.8449 7.6186.987.405.3687.8673.8334 1.0515 1.3806.2207.6913-.7695-.073-.9057-.167-.71-.4532-1.4182-.8334-2.2127-1.0946C12.8614.3873 8.8122 2.709 6.2945 6.315c-1.0516 1.5939-1.7367 3.2721-2.299 5.1174-.0614.2017-.1094.4647-.2207.6413-.1113.2036-.048-.5453-.048-.5702.0845-.7623.2438-1.4997.4414-2.237C5.3292 5.3375 7.897 2.0655 11.5891.5658zm4.9281 7.0587c0 1.6686-1.353 3.0224-3.0205 3.0224-1.6677 0-3.0186-1.3538-3.0186-3.0224 0-1.6687 1.351-3.0224 3.0186-3.0224 1.6676 0 3.0205 1.3518 3.0205 3.0224Z"/></svg>
|
||||
|
After Width: | Height: | Size: 933 B |
Loading…
Add table
Add a link
Reference in a new issue