Web Scraping API
Extract structured data from any website with a single API call. The scraper automatically follows links, extracts content, and returns structured JSON data.
Base URL: https://api.tafscraper.com
Start Scraping
POST /api/v1/scrape
Check Job Status
GET /api/v1/jobs/:job_id
Authentication
X-API-Key header
Authentication
All API requests require authentication using an API key. Include your key in the X-API-Key header.
X-API-Key: sk_your_api_key_hereGet your API key from the dashboard.
Credit System
- Free Credits — Can only be used for basic scraping (no context).
- Purchased Credits — Required for contextual/AI scraping.
- Each page scraped consumes 1 credit.
/api/v1/scrapeDescription
Starts a web scraping job for the specified URL. Returns cached results immediately if available, otherwise starts an asynchronous job.
Request Body
{
"url": "https://example.com",
"context": "Extract all email addresses",
"config": {
"max_pages": 10,
"max_depth": 2,
"delay_ms": 1000,
"follow_external_links": false,
"scrape_css_branding": false
}
}Parameters
url(required)The URL to scrape. Must be a valid HTTP/HTTPS URL.
context(optional)(Requires Purchased Credits)AI extraction prompt. Specify what data you want to extract.
config.max_pages(optional, default: 10)Maximum number of pages to scrape.
config.max_depth(optional, default: 2)Maximum depth to follow links.
config.delay_ms(optional, default: 1000)Delay between requests in milliseconds.
config.follow_external_links(optional, default: false)Whether to follow links to external domains.
config.scrape_css_branding(optional, default: false)Whether to extract branding. Each page then includes logo (URL, or an SVG data URI for inline SVG logos), favicon, and colors (primary, secondary, tertiary as #rrggbb). Linked stylesheets are fetched to detect colors.
Response (Job Started)
{
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "started",
"message": "Scraping job started for URL: https://example.com",
"cached": false,
"data": null
}Response (Cached)
{
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "completed",
"message": "Scraping completed",
"cached": true,
"data": [
{
"url": "https://example.com",
"title": "Example Domain",
"text_content": ["Welcome to example.com", "..."],
"links": ["https://example.com/about"],
"images": ["https://example.com/logo.png"]
}
]
}/api/v1/jobs/:job_idDescription
Check the status of a scraping job. Poll this endpoint until status is completed or failed.
Path Parameters
job_id(required)The UUID returned by the scrape endpoint.
Response
{
"job_id": "550e8400-e29b-41d4-a716-446655440000",
"status": "completed",
"message": "Job completed successfully. Scraped 15 pages",
"progress": {
"pages_scraped": 15,
"total_links_found": 120,
"current_url": null
},
"results": null,
"error": null
}Status Values
pending— Job is queuedrunning— Job is in progresscompleted— Job finished successfullyfailed— Job encountered an error
Code Examples
cURL
curl -X POST https://api.tafscraper.com/api/v1/scrape \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com",
"config": {
"max_depth": 2,
"max_pages": 10
}
}'JavaScript
const response = await fetch('https://api.tafscraper.com/api/v1/scrape', {
method: 'POST',
headers: {
'X-API-Key': 'YOUR_API_KEY',
'Content-Type': 'application/json',
},
body: JSON.stringify({
url: 'https://example.com',
config: {
max_depth: 2,
max_pages: 10
}
})
});
const data = await response.json();
if (data.cached) {
console.log('Scraped data:', data.data);
} else {
// Poll for job completion
const jobId = data.job_id;
let status = 'pending';
while (status !== 'completed' && status !== 'failed') {
await new Promise(r => setTimeout(r, 2000));
const statusRes = await fetch(`https://api.tafscraper.com/api/v1/jobs/${jobId}`, {
headers: { 'X-API-Key': 'YOUR_API_KEY' }
});
const statusData = await statusRes.json();
status = statusData.status;
console.log('Job status:', status);
}
}Python
import requests
import time
headers = {
'X-API-Key': 'YOUR_API_KEY',
'Content-Type': 'application/json'
}
# Start scraping
response = requests.post('https://api.tafscraper.com/api/v1/scrape',
headers=headers,
json={
'url': 'https://example.com',
'config': {
'max_depth': 2,
'max_pages': 10
}
}
)
data = response.json()
if data['cached']:
print('Scraped data:', data['data'])
else:
job_id = data['job_id']
# Poll for completion
while True:
status_res = requests.get(
f'https://api.tafscraper.com/api/v1/jobs/{job_id}',
headers=headers
)
status_data = status_res.json()
if status_data['status'] in ['completed', 'failed']:
print('Final status:', status_data['status'])
break
time.sleep(2)Error Responses
401 Unauthorized
{
"error": "Invalid API key",
"status": 401
}402 Payment Required
{
"error": "Insufficient credits",
"status": 402
}404 Not Found
{
"error": "Job 550e8400-e29b-41d4-a716-446655440000 not found",
"status": 404
}429 Too Many Requests
{
"error": "Rate limit exceeded",
"status": 429
}Rate Limits
| Endpoint | Limit |
|---|---|
/api/v1/scrape | 30 requests / minute |
/api/v1/jobs/:job_id | 30 requests / minute |