Web scraping & data extraction · Diffbot Technologies Corporation

Diffbot

Extraction API that uses machine-learning models to return structured data (articles, products, discussions) from any page without site-specific rules.

Diffbot is fundamentally an extraction model rather than a fetch or proxy service: instead of writing CSS selectors per site, you send it a URL and its computer-vision and NLP models classify the page type and return structured fields — article text and author, product name and price, discussion posts, and so on — automatically. It also offers a full web Crawl product for site-wide extraction and Knowledge Graph, a pre-built graph of entities built by crawling large parts of the public web. Everything runs as a hosted API billed in credits per successful extraction, with no self-hosted option. This makes Diffbot most useful when a team needs consistent structured output across many different, unfamiliar sites rather than deep control over one known site's markup.

At a glance

Vendor Diffbot Technologies Corporation
Pricing model Subscription
Free tier Yes
Deployment Cloud
Open source No
Best for Teams that need structured data from many unfamiliar sites without maintaining per-site extraction rules.

Pricing

Monthly plans include a credit allowance consumed per successful API call, with per-credit overage pricing beyond the plan limit; higher tiers add crawl capacity and seats.

Plan Price Notes
Free $0/month 10,000 credits/month, all four extraction APIs, 5 requests/min
Startup $299/month 250,000 credits/month, 5 requests/sec, overage $0.001/credit
Plus $899/month 1,000,000 credits/month, up to 25 active crawls, 3 user seats, overage $0.0009/credit
Enterprise custom custom credit allotment, 100+ active crawls, dedicated support

Prices read from the vendor's own page on September 21, 2026. Vendors change prices; check the source before you budget.

Features

  • Automatic page-type classification (article, product, discussion, event, image)
  • ML-based extraction with no per-site rules to maintain
  • Full-site Crawl product for bulk extraction
  • Knowledge Graph of entities built from web-scale crawling
  • Natural-language query interface over extracted/graph data
  • Visual and NLP models trained specifically for web content
  • REST API with per-credit billing

Integrations

Profile last reviewed September 21, 2026

Alternatives

Diffbot in the index now

Terms to know

Related guides