Skip to content

For AI agents

Building an AI agent, LLM-powered integration, or automated compliance tooling? Everything you need to wire iGregulator into a language-model workflow lives here.

Three machine-readable resources document the entire product. One fetch each — no crawling required.

  • llms.txt — structured index per the llmstxt.org spec. ~15 KB, complete product map with links into the longer docs.
  • llms-full.txt — every /docs/* page concatenated into one file. ~260 KB. One fetch buys full context.
  • OpenAPI 3.1 spec — authoritative API schema. Every endpoint, request/response shape, auth scheme, rate-limit description.

Every page on this site also emits <link rel="alternate"> pointers to all three URLs so a crawler hitting the landing can auto-discover them. The homepage additionally returns Link response headers (RFC 8288 / RFC 9727) pointing at the catalog, OpenAPI spec, docs, and server card — so headless agents find them without parsing HTML.

Standards-based entrypoints under /.well-known/ (and the apex), for agents that look there first:

  • /.well-known/mcp/server-card.json — MCP Server Card (SEP-1649): server info, transport endpoint, and the full tool list. (Legacy /.well-known/mcp.json is still served too.)
  • /.well-known/api-catalog — API catalog (RFC 9727) linking the OpenAPI spec, docs, llms.txt, and the /v1/health status endpoint.
  • /.well-known/agent-skills/index.json — Agent Skills discovery index (v0.2.0). Currently ships a verify-gambling-license skill (digest-pinned SKILL.md).
  • /auth.md — how to authenticate: bearer API keys for the REST API and MCP server, or OAuth 2.1 sign-in for MCP connectors at https://mcp.igregulator.io/mcp/account (its RFC 9728 metadata is on mcp.igregulator.io; the authorization server is app.igregulator.io).

Content usage is declared via Content-Signal in robots.txt — search, ai-input, and ai-train are all permitted.

Deliberate design choices that make integration cleaner for agents.

Every error response carries details.reason + details.suggestion. Agents branch on reason instead of free-form text and surface suggestion to the user / caller as-is.

{
"error": "domain is not a valid hostname",
"code": "invalid_query",
"details": {
"field": "domain",
"reason": "not_a_valid_hostname",
"suggestion": "Pass a bare hostname — no scheme, no path, no underscores. Example: 'paddypower.com' or 'www.bet365.com'."
}
}

The full reason vocabulary is stable; see error handling for the code + reason matrix — branch on code and reason, because one code (rate_limited, quota_exceeded, invalid_query) covers several causes.

/v1/check answers with a top-level verdict — licensed, licensed_provisional, licence_not_active, domain_not_listed, related_host_listed, name_match_only, not_found or generic_term — derived from status, domain_status, status_qualifier and confidence by fixed rules (confidence scoring), so an agent does not have to combine them itself. status: active with domain_status: delisted is domain_not_listed, not “licensed”; a revoked or suspended licence is licence_not_active even when the domain is de-listed too. Only licensed and licensed_provisional mean licensed now — treat every other value, including one you have not seen before, as “not confirmed as licensed”.

verdict_detail is one plain-English sentence to quote as it stands: it names the regulator, operator and licence (never our KH/TGC/IOM reference as the regulator’s number), dates the read that listed the domain and names any register behind the answer whose last read is past its freshness window, scopes a miss to the registers we cover, and never calls anything “unlicensed”. Branch on verdict; quote verdict_detail.

www.example.com and example.com are the same site: either finds the host the register lists (match.matched_domain says which). Any other subdomain is a different host. When the host you asked about is not listed but others on the same registrable domain are, the answer is never licensed: it is related_host_listed (or licence_not_active / domain_not_listed when that other host’s licence or listing is not in force), and related_hosts[] names those hosts, each with its operator. Say which hosts are listed and that this one is not; never carry their licence over to it.

Data-returning endpoints include a _meta envelope with provenance:

  • scraped_at — ISO-8601 timestamp when we last pulled this record.
  • source_url — exact regulator URL, or null if the row doesn’t map to one URL. For a licence this is the page that published its status (status_source_url) — an enforcement register or a revoked-licences list, say — which is often not the register the licence is listed in.
  • confidence_hint — authoritative (direct register dump), scraped (HTML / PDF), or derived (fuzzy match, not a direct lookup).
  • source_modified_at — reserved for the regulator’s own modification timestamp. Always null today: we don’t store it for any register yet.

On /v1/check, _meta also carries:

  • checked_at — when the answer was computed. (On a name match or a miss, scraped_at is that same request time — there is no record behind the answer.)
  • register — { jurisdiction, last_read_at, fresh, sla_hours } for the matched jurisdiction: when we last read its register and whether that is inside its freshness SLA (the same rule as /v1/health/coverage).
  • stale_jurisdictions — on a miss, a name-only match or a related-host answer, [{ code, last_read_at }] for every covered register past its SLA. A site listed there since our last read is not in our data yet: say so next to a “not found”.

And match carries what a citation needs without a second call: regulator_name, license_id (for GET /v1/licenses/{id}), license_reference_is_ours, status_source_url and status_observed_at (where this status was published, and the latest read that still said it), matched_domain (the host the register lists — www.x.com for x.com) and domain_last_listed_at (when a regulator source last listed it; null once de-listed).

An agent can say “verified via UKGC official register, scraped 6 h ago” with a real evidence trail, not a gloss. Which page can set which status in each jurisdiction, and what we keep as evidence, is on the methodology page.

verification_url — cite the regulator, not us

Section titled “verification_url — cite the regulator, not us”

Where the regulator runs its own per-domain verification page — Curaçao’s certificate portal, Tobique’s validation seal — match.verification_url on /v1/check links straight to it. That’s a primary-source citation an agent can hand to the user (“verify on the regulator’s own page”), and it’s how we keep the row fresh: every stored verification page is re-read on a rolling nightly pass (every Tobique seal each night, each Curaçao certificate every 2–3 days), and domain_status moves only on a clean read of the regulator’s own words. null for regulators with no such surface (UKGC, MGA, KH, AN, IOM). A Curaçao link stored on the legacy host cert.gcb.cw is served on cert.cga.cw.

Before you cite it, read verification_page_status — what that page printed about the licence at our latest read, verbatim (Active, Revoked, a Tobique seal’s VALID), with verification_page_read_at. The same pair is on every jurisdictions[] row and on each domains[] entry of /v1/operators/{slug}. If the word does not say the licence is in force, the page is not proof: say what it reads and when, and don’t offer it as confirmation. It changes no status on its own — a revocation comes only from a regulator publication (status_source_url).

A brand can be licensed by different legal entities in different jurisdictions at once. When a matched domain has more than one (operator, jurisdiction) pair, /v1/check carries a best-first jurisdictions[] array with per-link status, status_qualifier, domain_status, license_reference_is_ours, verification_url and register (when we last read that link’s register, and whether the read is fresh). Report all of them — the same domain can be active under one register and delisted under another, and naming only one is a wrong answer. See confidence scoring.

Public endpoint exposing per-jurisdiction scraper freshness. status: healthy vs degraded, age_hours, record_count per regulator. See endpoints. Useful for SLA dashboards and status pages.

Every endpoint has a stable operationId (e.g. checkDomain, checkDomainBatch, searchOperators, getOperator, getOperatorRegulatoryActions, listJurisdictions, getLicense, getLicenseHistory, checkCoverage). SDK generators and MCP servers use these as function names — no getV1CheckDomain slug noise.

A call made without a key to a public endpoint carries the policy in both a custom and the IETF-draft format. Parse either:

X-RateLimit-Policy: tier=public;limit=10;window=hour
RateLimit-Policy: "default";q=10;w=3600

The IETF draft (draft-ietf-httpapi-ratelimit-headers) is what Cloudflare, Kong, and similar gateways auto-parse; the custom format is human-readable for logs.

Keyed calls don’t send RateLimit-Policy. Their X-RateLimit-Limit / -Remaining / -Reset describe your monthly quota, and X-RateLimit-Policy appears only as tier=unlimited (no monthly cap) or tier=authenticated (a key on a public endpoint). See rate limits.

A field marked for removal is announced with standard headers (example shape):

Deprecation: @1790812800
Sunset: Mon, 01 Mar 2027 00:00:00 GMT
Link: <https://igregulator.io/docs/changelog/>; rel="deprecation"; type="text/html"

Deprecation is RFC 9745: a structured-field date (@ + Unix epoch seconds) saying when the field was deprecated, plus the rel="deprecation" link. Sunset is RFC 8594: an HTTP-date for the removal, minimum 90 days out. An agent caching request shapes can inspect Sunset before assuming stability. No fields are currently deprecated, so the API sends neither header today.

Live at mcp.igregulator.io. Streamable HTTP transport (current MCP spec — SSE used for streaming responses). Compatible with Claude Desktop, Cursor, Windsurf, Cline, and any other client that speaks MCP.

Tools exposed (lean output — a compact verdict, not the full REST payload):

  • check_domain — verify a licence by domain (supports as_of)
  • check_domain_batch — up to 100 domains in one call (KYB sweep)
  • search_operators — search the register by name
  • get_operator — full operator detail (supports as_of)
  • get_operator_regulatory_actions — enforcement history (fines, suspensions)
  • check_coverage — data freshness per jurisdiction
  • list_jurisdictions — all covered regulators
  • get_jurisdiction — single regulator metadata
  • get_license — single licence detail
  • get_license_history — status-change timeline

check_domain and every check_domain_batch row lead with the API’s verdict and verdict_detail — branch on the first, quote the second. On a no-match, check_domain also returns match_absence_reason + checked_jurisdictions — never collapse a miss into a bare “unlicensed”. On a match, verification_url (the regulator’s own page) and — for dual-licensed domains — a compact jurisdictions[] array pass through, so the verdict can be cited and never flattened to one register.

Auth = same API keys as the direct HTTP API — and optional: with no Authorization header, check_domain, search_operators (top 3), list_jurisdictions and check_coverage answer under the REST API’s public limit of 10 requests per hour per IP; the other six tools need a key. Each keyless tool call counts against the per-IP counter of the REST route it calls, from your IP: check_domain spends the same /v1/check allowance as a direct request (rate limits). Setup walkthrough + example prompts at /docs/mcp. Discovery via the MCP Server Card (SEP-1649) and the legacy /.well-known/mcp.json.

Distinct from the server above: the homepage registers WebMCP tools via navigator.modelContext, so an agent driving a browser can act without wiring up the HTTP/MCP integration at all. Two read-only tools, backed by the public API (10 req/IP/hour, no key):

  • check_gambling_license — verify by domain or licence number.
  • search_gambling_operators — search operators by name.

They’re feature-detected, so they simply don’t appear in browsers without the WebMCP API. Use the server-side MCP for production integrations; WebMCP is the zero-setup path for browser agents.

Payment-processor KYB flow:

  1. User submits a merchant application with a domain.
  2. Agent calls GET /v1/check?domain=X.
  3. Branch on verdict, and put verdict_detail in front of whoever decides:
    • licensed → approve. If a jurisdictions[] array is present, the domain is on more than one (operator, jurisdiction) licence and each entry has its own status and domain_status — read them all and report the full picture, not just match (confidence scoring).
    • licensed_provisional (status: active with status_qualifier: provisional_under_assessment) → licensed now, provisionally: a Curaçao licence past its stated term whose final assessment by the CGA is outstanding (the CGA keeps it in force until it decides). Approve if your policy accepts provisional licences, say so in the answer, and re-check — it can become final or end.
    • licence_not_active → do not auto-approve. verdict_detail names the status. revoked / suspended mean a regulator published an enforcement decision against that licence — cite match.status_source_url. expired, surrendered (the operator gave it up) and not_in_register (the register no longer lists it, and published no reason) mean “not licensed right now”, and unknown means wording we could not classify — none of them is a revocation, so route them to a human rather than rejecting with an enforcement-flavoured reason. upstream_status is the regulator’s own word; for since when a licence has been unlisted, fetch match.license_id from GET /v1/licenses/{id} (not_listed_since, last_listed_at).
    • domain_not_listed → do not auto-approve: the licence is active, but the regulator no longer lists this domain on it.
    • related_host_listed → manual review. This host is not listed; other hosts on the same registrable domain are (related_hosts[], each with its operator). The applicant’s site may be one of them under another name, or a different site altogether — ask, then check the host they confirm.
    • name_match_only → manual review. The domain is on no licence we read; a name matched an operator (confidence: medium is a close match, low a weak one), so status is that operator’s licence, not this site’s. Never reject on it — and never approve on it either.
    • generic_term → manual review. The domain root is a generic gambling label (casino.org, poker.com) and /v1/check refuses to guess which operator runs it. The site may well be licensed, just not identifiable from the label.
    • not_found → not in any covered register. Scope it to checked_jurisdictions (“not found in the registers of the 7 jurisdictions we cover”), never an unqualified “unlicensed”, and mention any _meta.stale_jurisdictions — verdict_detail already does both.
    • A value not listed here → treat it as “not confirmed”: manual review.

Portfolio-monitoring automation:

  1. Agent keeps a list of N operator slugs in its CRM / knowledge base.
  2. Daily cron: iterate the list, call GET /v1/operators/:slug (licences
    • domains) and GET /v1/operators/:slug/regulatory-actions (enforcement history — not included in the operator detail). An empty list there means no action is linked to that operator, not a clean record: many published actions aren’t matched to an operator (UKGC 107 of 107 linked, MGA 2 of 160, CW 0 of 18 on 2026-09-28). The response says so itself: _meta.note (quote it on an empty list) and _meta.sources_read, the regulator publications we read.
  3. Diff against previous day’s snapshot. Detect status changes, regulatory actions, expiry windows. (Or let webhooks on a watchlist push them to you.)
  4. Alert compliance team on anomalies.

Risk-scoring for gambling-adjacent domains:

  1. Agent receives an unknown gambling-related domain.
  2. Calls GET /v1/check?domain=X.
  3. When the verdict ties the domain to a licence (licensed, licensed_provisional, licence_not_active, domain_not_listed), fetches the operator’s enforcement history from GET /v1/operators/{slug}/regulatory-actions and combines verdict + any regulatory actions into a composite risk score. A name_match_only operator is not tied to the domain: its record says nothing about this site.
  4. Score feeds the downstream decision (list / delist / require extra verification).

Auditing a whole merchant book or affiliate list in one shot:

  1. Collect the domains (chunks of 100).
  2. POST /v1/check/batch with { "domains": [...] } — one tool-call per 100 instead of one per domain.
  3. Iterate results: each row carries verdict + verdict_detail + match + confidence + match_absence_reason (same semantics as the single check), plus related_hosts[] when the host isn’t stored but others on its registrable domain are, and query.input, the string you sent. A row has no jurisdictions[] and no as_of — check a dual-licensed domain singly for its other links. An entry we can’t look up comes back with error and no verdict — report it as not checked, never as not found — and doesn’t fail the batch.
  4. checked_jurisdictions and _meta.stale_jurisdictions are returned once at the top — use them to scope every “not found” verdict. See Batch domain check.

“Was this merchant licensed at the time of the transaction?”:

  1. GET /v1/check?domain=X&as_of=2026-03-01 (or /v1/licenses/{id}?as_of=). as_of is a YYYY-MM-DD date (end of that day, UTC; today’s date means now) or an ISO-8601 datetime with a UTC offset — anything else, an impossible date or a future one is a 400.

  2. verdict still describes today; verdict_detail leads with the answer for your date (“On 2026-03-01 we were not yet tracking … licence … (we first recorded it on …), so its status then is unknown. Today: …”). On /v1/check the as_of object names the licence it answered for (license_id, license_number, operator, scope: "licence"): today’s licence for that domain. We keep no record of when a domain was linked to a licence (link_note says so), so this is that licence’s status on the date — not proof of who ran the site then.

  3. Read the as_of object, and honour knowledge:

    • observed → status_as_of is the real status then; established_by shows when it was last confirmed relative to your date.
    • before_tracking → the date predates our observation window. Do not assert a status — tell the user we weren’t watching before tracking_since. This is the difference between a defensible answer and a fabricated one.
    • corrected → what we recorded for that date was later withdrawn. status_as_of is null; correction.withdrawn_status is what we used to say, never the status. Treat it like before_tracking.
    • no_such_license → we hold no history for that licence. null status.
    • no_license_resolved (/v1/check only) → the match didn’t resolve to a specific licence, so there is nothing to time-travel. null status.

    Only observed carries a status. See Point-in-time lookups.

  1. Fetch llms-full.txt for complete docs context in a single request.
  2. Review the OpenAPI spec.
  3. Try endpoints interactively in the playground.
  4. Create a free account for an API key and MCP server access — free for founding members.

Questions, integration help, feedback — founder@igregulator.io.