Guides · 28
Guides to public data APIs and scraping
Each guide explains one public data source in depth: whether an official API exists, what it limits or blocks, and a runnable way to get the data out. Every guide links to the free tool on this site that builds the matching query.
Updated · By Omar Eldeeb
All guides
ClinicalTrials.gov API to CSV: Export Trials by Phase in PythonExport ClinicalTrials.gov search results to CSV with the free v2 API: format=csv, the x-next-page-token header, and the phase filter that actually works.Cointelegraph API: Get Crypto News as JSON Without a KeyCointelegraph has no public news API, but its site runs on a keyless GraphQL endpoint and RSS feeds. Pull articles, search results and full text as JSON.Price per sqm by Compound in New Cairo: Compute It with PythonStop trusting broker blogs: compute New Cairo apartment price per square meter by compound from live Property Finder Egypt listings, split ready vs off-plan.FPL API Price Change Data in Python: The Official 2026/27 FieldsRead FPL's own price-change predictions from the free bootstrap-static API in Python: price_change_percent, projections, likelihood, and the ±100 rule.LinkedIn X-Ray Search for Recently Updated Profiles (Past Week)Add Google's past-week filter to a LinkedIn X-ray search to surface recently re-indexed profiles, then dedupe by slug so each run shows only new people.Measure AI Share of Voice in Python: ChatGPT, Gemini, PerplexityMeasure your brand's AI share of voice in Python: one prompt set, several LLMs, mention counting that ignores 'I'm not familiar with X' answers.How to Check a DEX Pool's Sandwich Attack Rate (Python + API)Measure how often a Uniswap, Raydium or Aerodrome pool gets sandwiched: pull per-hour sandwichRate and MEV fee data from Codex getBars in Python.Track Peptide & GLP-1 Mentions on YouTube with Python (No API Key)Count YouTube videos mentioning semaglutide, tirzepatide or BPC-157 without an API key: parse ytInitialData, map brand names to compounds, export to CSV.Scrape Monthly vs Annual SaaS Prices From Pricing-Page JSON-LDGet both sides of a SaaS pricing page's monthly/annual toggle without a browser: read schema.org Offer JSON-LD in Python, then diff it on a schedule.Shopify products.json Pagination: 250 Limit and the 25,000 CapPaginate any Shopify store's public /products.json: limit=250, page numbers, the empty-page stop signal, and the HTTP 400 you hit at 25,000 products.Threads Keyword Search Without App Review: Monitor in PythonThreads' keyword_search API only searches your own posts until Meta approves your app. Monitor public Threads keywords in Python without app review.PERM Disclosure Data in Python: Find Green Card Sponsors by CityUse the DOL's free PERM disclosure file to list green card sponsors by city and occupation in Python: where the file lives, key columns, and the traps.How to Scrape YouTube Shorts Data (Exact View, Like & Comment Counts)A practical guide to scraping YouTube Shorts data: why the official API falls short, how the internal API works, and a runnable example with exact counts.How to Extract Dubai Used Car Listings Data from Gulf MarketplacesA practical, honest guide to extracting Dubai used car listings data from DubiCars, YallaMotor, and Dubizzle Motors using JSON-LD and embedded __NEXT_DATA__.How to Build an H-1B Salary Database by Employer (with Python)A practical guide to querying H-1B salary data by employer and role, using the authoritative DOL OFLC LCA disclosure files with runnable Python.How to Extract Saudi Arabia Property Data from 4 Listing PortalsA practical guide to collecting Saudi Arabia property data from the major portals, including the REGA license trick that dedupes listings across platforms.How to Scrape a Telegram Channel Without Login (No API Key)A verified, runnable guide to scraping public Telegram channels without login, an API key or a phone number, using the t.me/s/ preview and plain HTTP.Understat xG Data Export: Pull Expected Goals with Python + CSVA practical guide to Understat xG data export: how the data is embedded in the page, a Python scraper that decodes it, and league and player xG to CSV.Facebook Ad Library Scraper: API Limits and the Real ApproachWhy the official Meta Ad Library API only covers political ads, and how to build a Facebook ad library scraper for commercial competitor ads instead.How to Build a LinkedIn Profile Scraper: The Honest Technical GuideA practical, honest guide to building a LinkedIn profile scraper: why there's no public API, how public pages embed JSON-LD, and the legal reality.The SEC EDGAR API: A Practical Guide to Free Filing Data in PythonA hands-on guide to the free SEC EDGAR API: the required User-Agent header, ticker-to-CIK mapping, XBRL financials, and full-text search in Python.How to Build a Threads Scraper for Meta Profiles and PostsA practical, honest guide to building a Threads scraper for Meta's threads.com — what loads cookie-free, what doesn't, and a runnable example.The TikTok Ad Library API: A Developer's Guide to the DSA LibraryA practical, accurate guide to the TikTok ad library API: what the DSA Commercial Content Library covers, the official researcher path, and how to query it.App Store Top Charts API: Free, Key-Free, and CORS-OpenPull Apple App Store top charts straight from the browser with the legacy iTunes RSS feed — no key, CORS-open, runnable fetch() + curl included.The Hacker News Search API: Free, No-Key, and Surprisingly PowerfulA practical guide to the free, no-key Hacker News search API (hn.algolia.com): tags, numeric filters, relevance vs. date sorting, and runnable code.Read Company Hiring Signals From Public Job Board APIs (with code)Open job postings leak strategy. Learn to read company hiring signals from the public, no-auth Greenhouse API with a runnable JS classifier.How to Export Google Patents to CSV (Honest Guide to Every Real Path)A precise, no-hype guide to exporting Google Patents search results to CSV — the built-in 1,000-row cap, the BigQuery bulk path, and why browser scraping fails.How to Scrape Reddit Without the API (After the 2023 Price Changes)A precise, honest guide to scraping Reddit without the API: what still works in 2026, the login-wall and 403/CORS traps, and old.reddit.com.