API Documentation

Web Scraping API

Extract structured data from any website with a single API call. The scraper automatically follows links, extracts content, and returns structured JSON data.

Base URL: https://api.tafscraper.com

Authentication

All API requests require authentication using an API key. Include your key in the X-API-Key header.

X-API-Key: sk_your_api_key_here

Get your API key from the dashboard.

Credit System

  • Free Credits — Can only be used for basic scraping (no context).
  • Purchased Credits — Required for contextual/AI scraping.
  • Each page scraped consumes 1 credit.
POST/api/v1/scrape

Description

Starts a web scraping job for the specified URL. Returns cached results immediately if available, otherwise starts an asynchronous job.

Request Body

{
  "url": "https://example.com",
  "context": "Extract all email addresses",
  "config": {
    "max_pages": 10,
    "max_depth": 2,
    "delay_ms": 1000,
    "follow_external_links": false,
    "scrape_css_branding": false
  }
}

Parameters

url(required)

The URL to scrape. Must be a valid HTTP/HTTPS URL.

context(optional)(Requires Purchased Credits)

AI extraction prompt. Specify what data you want to extract.

config.max_pages(optional, default: 10)

Maximum number of pages to scrape.

config.max_depth(optional, default: 2)

Maximum depth to follow links.

config.delay_ms(optional, default: 1000)

Delay between requests in milliseconds.

config.follow_external_links(optional, default: false)

Whether to follow links to external domains.

config.scrape_css_branding(optional, default: false)

Whether to extract branding. Each page then includes logo (URL, or an SVG data URI for inline SVG logos), favicon, and colors (primary, secondary, tertiary as #rrggbb). Linked stylesheets are fetched to detect colors.

Response (Job Started)

{
  "job_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "started",
  "message": "Scraping job started for URL: https://example.com",
  "cached": false,
  "data": null
}

Response (Cached)

{
  "job_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "completed",
  "message": "Scraping completed",
  "cached": true,
  "data": [
    {
      "url": "https://example.com",
      "title": "Example Domain",
      "text_content": ["Welcome to example.com", "..."],
      "links": ["https://example.com/about"],
      "images": ["https://example.com/logo.png"]
    }
  ]
}
GET/api/v1/jobs/:job_id

Description

Check the status of a scraping job. Poll this endpoint until status is completed or failed.

Path Parameters

job_id(required)

The UUID returned by the scrape endpoint.

Response

{
  "job_id": "550e8400-e29b-41d4-a716-446655440000",
  "status": "completed",
  "message": "Job completed successfully. Scraped 15 pages",
  "progress": {
    "pages_scraped": 15,
    "total_links_found": 120,
    "current_url": null
  },
  "results": null,
  "error": null
}

Status Values

  • pending — Job is queued
  • running — Job is in progress
  • completed — Job finished successfully
  • failed — Job encountered an error

Code Examples

cURL

curl -X POST https://api.tafscraper.com/api/v1/scrape \
  -H "X-API-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com",
    "config": {
      "max_depth": 2,
      "max_pages": 10
    }
  }'

JavaScript

const response = await fetch('https://api.tafscraper.com/api/v1/scrape', {
  method: 'POST',
  headers: {
    'X-API-Key': 'YOUR_API_KEY',
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    url: 'https://example.com',
    config: {
      max_depth: 2,
      max_pages: 10
    }
  })
});

const data = await response.json();

if (data.cached) {
  console.log('Scraped data:', data.data);
} else {
  // Poll for job completion
  const jobId = data.job_id;
  let status = 'pending';
  
  while (status !== 'completed' && status !== 'failed') {
    await new Promise(r => setTimeout(r, 2000));
    const statusRes = await fetch(`https://api.tafscraper.com/api/v1/jobs/${jobId}`, {
      headers: { 'X-API-Key': 'YOUR_API_KEY' }
    });
    const statusData = await statusRes.json();
    status = statusData.status;
    console.log('Job status:', status);
  }
}

Python

import requests
import time

headers = {
    'X-API-Key': 'YOUR_API_KEY',
    'Content-Type': 'application/json'
}

# Start scraping
response = requests.post('https://api.tafscraper.com/api/v1/scrape', 
    headers=headers,
    json={
        'url': 'https://example.com',
        'config': {
            'max_depth': 2,
            'max_pages': 10
        }
    }
)

data = response.json()

if data['cached']:
    print('Scraped data:', data['data'])
else:
    job_id = data['job_id']
    
    # Poll for completion
    while True:
        status_res = requests.get(
            f'https://api.tafscraper.com/api/v1/jobs/{job_id}',
            headers=headers
        )
        status_data = status_res.json()
        
        if status_data['status'] in ['completed', 'failed']:
            print('Final status:', status_data['status'])
            break
        
        time.sleep(2)

Error Responses

401 Unauthorized

{
  "error": "Invalid API key",
  "status": 401
}

402 Payment Required

{
  "error": "Insufficient credits",
  "status": 402
}

404 Not Found

{
  "error": "Job 550e8400-e29b-41d4-a716-446655440000 not found",
  "status": 404
}

429 Too Many Requests

{
  "error": "Rate limit exceeded",
  "status": 429
}

Rate Limits

EndpointLimit
/api/v1/scrape30 requests / minute
/api/v1/jobs/:job_id30 requests / minute
Ask TafScraper's AI