Web scraping & data extraction · Diffbot Technologies Corporation
Diffbot
Extraction API that uses machine-learning models to return structured data (articles, products, discussions) from any page without site-specific rules.
Diffbot is fundamentally an extraction model rather than a fetch or proxy service: instead of writing CSS selectors per site, you send it a URL and its computer-vision and NLP models classify the page type and return structured fields — article text and author, product name and price, discussion posts, and so on — automatically. It also offers a full web Crawl product for site-wide extraction and Knowledge Graph, a pre-built graph of entities built by crawling large parts of the public web. Everything runs as a hosted API billed in credits per successful extraction, with no self-hosted option. This makes Diffbot most useful when a team needs consistent structured output across many different, unfamiliar sites rather than deep control over one known site's markup.
At a glance
| Vendor | Diffbot Technologies Corporation |
|---|---|
| Pricing model | Subscription |
| Free tier | Yes |
| Deployment | Cloud |
| Open source | No |
| Best for | Teams that need structured data from many unfamiliar sites without maintaining per-site extraction rules. |
Pricing
Monthly plans include a credit allowance consumed per successful API call, with per-credit overage pricing beyond the plan limit; higher tiers add crawl capacity and seats.
| Plan | Price | Notes |
|---|---|---|
| Free | $0/month | 10,000 credits/month, all four extraction APIs, 5 requests/min |
| Startup | $299/month | 250,000 credits/month, 5 requests/sec, overage $0.001/credit |
| Plus | $899/month | 1,000,000 credits/month, up to 25 active crawls, 3 user seats, overage $0.0009/credit |
| Enterprise | custom | custom credit allotment, 100+ active crawls, dedicated support |
Prices read from the vendor's own page on September 21, 2026. Vendors change prices; check the source before you budget.
Features
- Automatic page-type classification (article, product, discussion, event, image)
- ML-based extraction with no per-site rules to maintain
- Full-site Crawl product for bulk extraction
- Knowledge Graph of entities built from web-scale crawling
- Natural-language query interface over extracted/graph data
- Visual and NLP models trained specifically for web content
- REST API with per-credit billing
Integrations
Profile last reviewed September 21, 2026