SnapplAI – Apply to Jobs in a Snap SnapplAI, an open-source job-hunting pipeline created by TDK-99, scrapes LinkedIn listings, uses Google's Gemini 3.5 Flash Lite AI to summarize and score each job against a user's CV, and emails only the top matches. The pipeline runs in four automated stepsβ€”scrape, summarize, analyze, deliverβ€”and is available on GitHub with setup via Docker or GitHub Actions. Stop refreshing LinkedIn. This pipeline scrapes new job listings based on your settings, uses AI agents to summarize each one and score it against your CV, then delivers only the best matches straight to your inbox πŸ“¬ β€” so you're always first to apply πŸš€ The Problem -the-problem How It Works -how-it-works Tech Stack -tech-stack Pipeline Architecture -pipeline-architecture AI Output Fields -ai-output-fields Setup -setup Project Structure -project-structure Roadmap v2 -roadmap-v2 Contributing -contributing License -license Job hunting on LinkedIn is a full-time job in itself. New listings appear daily, most are irrelevant, and by the time you spot a good one, 200 people have already applied. SnapplAI flips the game: it runs on a schedule, scrapes fresh listings, lets AI read and score every single one against your CV, and emails you only the top matches β€” before the crowd even sees them. The pipeline runs in 4 sequential steps, fully automated: 1. Scrape β†’ job scraper pulls fresh listings from LinkedIn based on your search settings role, location, filters using python-jobspy. 2. Summarize β†’ agentic summarize sends each job description to Gemini, which extracts structured fields title, seniority, skills, salary, etc. as clean JSON. 3. Analyze β†’ agentic analyze reads your CV and scores each listing on how well it matches your profile. Chain-of-thought enforced: the model writes analysis before score in the JSON schema, so reasoning comes before judgment. 4. Deliver β†’ send email builds an email with the top-scored jobs and sends it to your inbox via SMTP. Key principle: AI reads and evaluates. Python orchestrates and delivers. No frameworks, no agents-calling-agents β€” just a clean data pipeline with LLM calls where they matter. | Component | Technology | |---|---| | LLM | Google GenAI SDK β€” gemini-3.5-flash-lite | | Scraping | python-jobspy LinkedIn | | Data | pandas, PyPDF / PyMuPDF | | Parsing | BeautifulSoup4 | | smtplib SMTP | | | Config | python-dotenv | The entire pipeline operates on a single pandas DataFrame that gets enriched at each step. No intermediate files, no database β€” everything flows through memory. Each job in the email is ranked by match score and includes company, role, work mode, a one-line AI summary explaining why it matched or didn't , and a direct apply link to the LinkedIn listing. - Get a free API key from Google AI Studio https://aistudio.google.com/apikey - Generate a Gmail App Password https://myaccount.google.com/apppasswords - Place your CV PDF in your cv config/ - Configure search settings: use file config.txt to create your file config.env filter docs https://github.com/Bunsly/JobSpy - Create your .env from the template: cp example env.txt .env git clone https://github.com/TDK-99/SnapplAI.git && cd SnapplAI pip install -r requirements.txt complete setup steps above python main.py git clone https://github.com/TDK-99/SnapplAI.git && cd SnapplAI complete setup steps above docker build -t snapplai . docker run --env-file .env snapplai - Fork this repo or create a private copy private-copy - Complete setup steps 1-4 above in your fork - Edit your settings in .github/workflows/snapplai.yml under the env: block - Add credentials as repository secrets Settings β†’ Secrets β†’ Actions : GOOGLE API KEY , GMAIL USER , GMAIL APP PASSWORD Actions tab β†’ enable workflows β†’ Run workflow SnapplAI/ β”œβ”€β”€ main.py Entry point β€” runs the 4-step pipeline β”œβ”€β”€ src/ β”‚ β”œβ”€β”€ daily scraper.py LinkedIn scraping with python-jobspy β”‚ β”œβ”€β”€ ai agents.py Gemini calls: summarize + analyze β”‚ └── smtp.py Email builder and SMTP sender β”œβ”€β”€ your cv config/ β”‚ β”œβ”€β”€ .gitkeep Keeps folder tracked in git β”‚ β”œβ”€β”€ file config.env Your settings role, location, filters β”‚ β”œβ”€β”€ file config.txt Additional config parameters β”‚ └── Your CV.pdf Your CV goes here PDF β”œβ”€β”€ .github/ β”‚ └── workflows/ β”‚ └── snapplai.yml GitHub Actions workflow scheduled + manual β”œβ”€β”€ Dockerfile Run anywhere with Docker β”œβ”€β”€ .env API keys and SMTP credentials git-ignored β”œβ”€β”€ example env.txt Template for .env variables β”œβ”€β”€ requirements.txt Dependencies β”œβ”€β”€ LICENSE MIT └── README.md Multi-country scraping β€” search across 2+ countries in a single run custom feature, not supported by python-jobspy out of the box Excel/DB deduplication β€” persistent storage to compare runs and filter out already-seen listings, so you never score the same job twice Scoring calibration β€” benchmark AI scores against known good/bad matches to improve match quality Output redesign β€” better visual formatting for the email report job cards, readability, direct links Contributions are welcome β€” bug fixes, new features, or docs improvements. Issues β€” Report bugs or suggest features Pull Requests β€” Fork, build, submit MIT β€” see LICENSE /TDK-99/SnapplAI/blob/main/LICENSE