MH MediaHarvester Web Context API

Web Extraction

From a Public Page to Reliable Web Context: Lessons from ION replaced homegrown brand scraping with a web context API in an afternoon

A MediaHarvester implementation guide prompted by ion replaced homegrown brand scraping with a web context api in an afternoon, with a tested API workflow, visual map and practical deployment decisions.

Web Extraction workflow
Public URL Engine Router Quality Score Clean Markdown Application

Why this subject deserves an implementation guide

The starting question for this article came from a public Context.dev post titled "ION replaced homegrown brand scraping with Context.dev in an afternoon". Rather than reproduce that article, this guide asks what the same product problem looks like inside MediaHarvester and what can be verified in the running application.

A page returning HTTP 200 is not the same thing as usable content. Navigation chrome, JavaScript placeholders, consent overlays and blocked responses can all produce text that looks fetched but should never be fed blindly to an application.

The MediaHarvester approach

MediaHarvester starts with an accessible fetch path, records the engine trace, scores the extracted result and exposes explicit blocked or verification states. The consumer can request clean markdown for an LLM, raw HTML for auditing, or a crawl when one URL is not enough.

The primary surface for this workflow is `GET /v1/web/scrape/markdown`. It can be tried from the API playground and integrated through the local API key, SDK, CLI or MCP layer.

Workflow map

Public URL Engine Router Quality Score Clean Markdown Application
A knowledge-base team is importing product documentation before a support assistant answers customer questions.

A realistic application scenario

A knowledge-base team is importing product documentation before a support assistant answers customer questions.

The workflow begins with a narrowly scoped public or authorized source, records the endpoint output and makes the result reviewable before it becomes visible to users or informs an automated decision.

Try the capability

GET /v1/web/scrape/markdown Local API key: mh-localhost-dev-key
GET /v1/web/scrape/markdown?url=https://web-tasarimci.com&maxAgeMs=0

Implementation choices that matter

In practice, teams begin with a small page budget and maxAgeMs set deliberately. Fresh mode is useful during testing; cached mode is a better default for stable product documentation. The same response metadata makes retries observable instead of mysterious.

This matters because a production feature is judged less by a perfect demo result than by how it behaves when an asset is missing, a source changes, a response is cached or a request is not allowed.

How to measure whether it works

Measure usable pages, low-quality results, refresh latency and citation coverage. A successful ingestion pipeline is not the one that fetched the most URLs; it is the one whose retrieved context remains accurate and explainable.

The app should retain enough source and request metadata to debug poor results while applying appropriate retention and access policies for customer data.

A responsible next step

Run the included endpoint against a website you control or are authorized to process, inspect the response in Visual and JSON modes, then decide which fields deserve automation and which deserve human approval.

MediaHarvester deliberately treats blocked, verification-required, login-required, robots-disallowed and permission-required outcomes as information, not obstacles to be bypassed.

FAQ

Questions teams ask before implementing this workflow

What does this web extraction workflow return?

It uses GET /v1/web/scrape/markdown and related MediaHarvester surfaces to return structured context together with metadata appropriate to the workflow.

Can I test this locally?

Yes. Run the local service at http://127.0.0.1:8013 and send X-API-Key: mh-localhost-dev-key to protected API routes.

Does it work with private or blocked pages?

The platform is designed for publicly accessible or authorized sources. Verification, login, permission and robots restrictions are reported rather than bypassed.

Can this be automated?

The same API surfaces are available through CLI, Python and TypeScript SDKs, MCP tools and starter no-code integration templates.

How do I keep the result current?

Use cache freshness controls such as maxAgeMs where exposed, and schedule refreshes in a production worker only as frequently as the business case needs.