跳到主要内容
知仓学习社ZHICANG

yc-jobs-scraper

Scrape daily job listings from YCombinator's Workatastartup platform without duplicates. Use this skill when asked to scrape YC jobs, update the YC …

写文件联网无严重或高危命中Varnan-Tech/opendirectory

它会碰到什么

扫了多少7 个文本文件,69 KB
它会碰到什么写文件联网
命中总数22 处
命中统计严重 0 · 高 0 · 中 1 · 低 21

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

YC Jobs Scraper

This skill provides a robust architecture for scraping jobs from YCombinator and workatastartup.com. It is designed to run automatically, bypass login bottlenecks, and maintain state to never scrape duplicate jobs.

Architecture

The scraper uses a hybrid approach to maximize reliability and minimize bot detection:

  1. Authentication: scripts/auth.js uses Playwright to let a human log in once and saves the session to scripts/state.json.
  2. Database: scripts/db.js uses better-sqlite3 to manage scripts/jobs.db. It tracks every company_slug and job_id ever seen.
  3. Primary Extraction: scripts/scraper.js loads state.json, visits YC query URLs, and extracts company slugs from the hidden Inertia.js data-page JSON payload.
  4. Job Extraction (JSON): It then visits the authenticated company pages (/companies/[slug]) to extract jobs from the backend JSON payload to ensure we get the real job_id for accurate deduplication.
  5. Job Extraction (Fallback): If the JSON extraction fails, it falls back to parsing public HTML job cards from ycombinator.com/companies/[slug]/jobs.

Workflows

1. First-Time Setup

If this is the first time running the scraper in an environment, or if node_modules is missing:

cd @path/scripts
npm install
npx playwright install

2. Authentication (Manual Step)

If scripts/state.json is missing or expired, the scraper will fail. You must instruct the human user to run the authentication script manually:

cd @path/scripts
node auth.js

Tell the user a browser will open, and they must log in. Playwright will automatically save the cookies/tokens to state.json.

3. Running the Daily Scraper

To scrape for new companies and jobs:

cd @path/scripts
node scraper.js

This script will output exactly how many new companies and new jobs were found. Because of jobs.db, running it multiple times consecutively will result in 0 new jobs found.

4. Querying the Database

If you need to analyze the scraped data or view the companies/jobs, you can query scripts/jobs.db directly using better-sqlite3.

Example: Count Companies

cd @path/scripts
node -e "const db = require('better-sqlite3')('jobs.db'); console.log('Companies:', db.prepare('SELECT COUNT(*) as count FROM companies').get().count);"

Example: View Recent Jobs

cd @path/scripts
node -e "const db = require('better-sqlite3')('jobs.db'); const jobs = db.prepare('SELECT title, company_slug, location FROM jobs ORDER BY created_at DESC LIMIT 5').all(); console.table(jobs);"

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。