MH MediaHarvester Web Context API

Images

Scrape Images

Extract page images, responsive sources, posters and inline SVGs.

GET /v1/web/scrape/images

What you get

Developer-ready output for Scrape Images

01 img and picture sources
02 CSS background candidates
03 Inline SVG and data URI handling
04 Dimensions, MIME and color enrichment

How it works

Four steps from request to usable JSON

1

Load the page

2

Collect visible and metadata image sources

3

Normalize absolute URLs

4

Return enriched image records

API response

Fields developers usually wire first

Extract images, responsive sources, CSS backgrounds, inline SVGs and video posters from any URL.

images[].src images[].alt images[].mimeType images[].classification
curl.exe -H "X-API-Key: mh-localhost-dev-key" "http://127.0.0.1:8013/v1/web/scrape/images"

Developer FAQ

Common questions about Scrape Images

12 answers
Does it find lazy-loaded images?

Yes, it checks common lazy attributes such as data-src, data-original and responsive srcset values in addition to normal img tags.

Can I classify logos versus photos?

Use enrichment to label common asset types such as logo, icon, graphic, pattern, product image or photography.

Does it download the image bytes?

The data API returns image URLs and metadata. Use the media scraper workspace when you want to download actual image files into a bundle.

What endpoint should I call for Scrape Images?

Use GET /v1/web/scrape/images. Local development accepts the X-API-Key header with mh-localhost-dev-key, and production clients can use the same shape with their own key.

Can I call it from CLI, SDK, MCP and no-code tools?

Yes. The same endpoint can be called with curl, the TypeScript SDK, Python SDK, MCP server, Zapier, Make, n8n, Google Sheets or Excel workflows where appropriate.

How does caching with maxAgeMs work?

When maxAgeMs is supported, a cached response can be reused if it is younger than the requested freshness window. Set maxAgeMs to 0 when you need a fresh scrape or extraction.

What happens when a page is blocked or requires login?

The API returns clear states such as blocked, verification_required, login_required, robots_disallowed or permission_required. It does not promise to bypass access controls.

Which response fields should I store?

Most integrations store images[].src, images[].alt, images[].mimeType, images[].classification, plus source URL, status, confidence and timestamp fields when present. Keep source metadata so users can audit values later.

Can I batch this endpoint?

For local or small workflows, loop over a list of URLs or domains. For larger jobs, use crawl, batch workflow endpoints or a queue so failures and retries are tracked cleanly.

How should I handle errors in production?

Check success, HTTP status, error.type and retryable flags. Retry timeouts and temporary network failures, but do not retry robots, login or permission errors without changing input or authorization.

Does the API work with JavaScript-heavy sites?

The engine router starts with fast fetching and escalates when content quality is low. Browser-style rendering is used where available, while access-control barriers are reported explicitly.

Is this safe to use with customer data?

Send only domains, URLs or fields your workflow needs. Use retention controls for generated artifacts and avoid storing unnecessary raw HTML, screenshots or extracted personal data.

Related data APIs

Build the next step in the workflow