For agents calling the decoder

Every failure says whether to try again.

A sentence tells a person what went wrong. An agent needs one more thing: should it retry, change the request, or stop? Every error from /api/decode-url, the MCP tools, the legibility check and the extractor carries a code from the list below and a retry verdict beside the sentence. Existing readers of error keep working: it is still the sentence.

What a refusal looks like

A real one. Google's robots.txt disallows /search for every crawler, so we do not read it:

{
  "error": "Could not read https://www.google.com/search?q=design: robots.txt disallows CSSCremeBot on /search (Disallow: /search)",
  "code": "ROBOTS_DENIED",
  "retry": "never",
  "status": 403,
  "hint": "The site owner asked crawlers not to read this path, and we obey that everywhere, including the /decode workspace. Nothing on our side will change the answer.",
  "robots": {
    "checked": true,
    "allowed": false,
    "agent": "*",
    "rule": "Disallow: /search",
    "source": "https://www.google.com/robots.txt",
    "status": 200
  },
  "errors": "https://csscreme.com/ai/errors"
}

Over MCP the same failure is a tool error whose text leads with the code and the verdict, so the first token is enough to branch on. The pair is also in _meta["com.csscreme/error"] as data.

[ROBOTS_DENIED] retry: never. Could not read https://www.google.com/search?q=design: robots.txt disallows CSSCremeBot on /search (Disallow: /search). …

The three retry words

never
The same request will fail the same way, and so will any rewording of it. Stop.
later
The same request may succeed after a wait. When we know how long, Retry-After says so.
with-changes
This request will keep failing. A corrected one (another URL, a valid parameter) may not.

Every code

CodeHTTPRetryWhat it means
MISSING_URL 400 with-changes No url was given.
INVALID_URL 400 with-changes Not an absolute http(s) URL on a standard port.
BLOCKED_ADDRESS 400 never The host is, or redirects to, a private, loopback, link-local or cloud-metadata address. Refused before anything is fetched.
INVALID_PARAM 400 with-changes A parameter other than the URL is malformed or unknown: a format name, an expect colour, a page count.
METHOD_NOT_ALLOWED 405 with-changes The endpoint answers a different HTTP method.
RATE_LIMITED 429 later Too many requests from one address in a minute. Retry-After says how long to wait.
ROBOTS_DENIED 403 never The site's robots.txt disallows CSSCremeBot on this path, and we obey it. Retrying will not change that, and neither will another CSS Crème tool: they all read through the same check.
ROBOTS_UNREACHABLE 503 later The site's robots.txt answered with a server error or could not be reached. RFC 9309 says to treat that as a full disallow until it answers.
TARGET_HTTP_ERROR 502 with-changes The site answered with an HTTP error status. The target status is in `targetStatus`; a 5xx there may clear later, a 4xx will not.
TIMEOUT 504 later The page or its stylesheets did not arrive within the time budget, or the host did not accept a connection in time.
HOST_NOT_FOUND 400 with-changes The host name does not resolve in DNS: a typo, a domain that does not exist, or one that has lapsed. Nothing was fetched.
TLS_FAILED 502 with-changes The site's HTTPS handshake failed. Many parked and old domains only answer plain http; if that is the address, pass the http:// URL.
FETCH_FAILED 502 later The connection failed before an answer arrived: DNS, TLS, a reset or a refused connection.
NOT_THE_SITE 422 never The page that answered is not the site: a bot challenge, a parked or for-sale domain, a host's default page, or a suspended account. `identity` names which, and the evidence. A challenge is never solved, so its verdict is never; a parked domain's is with-changes (decode the address the site really lives at).
BACKOFF 503 later This site failed several reads in a row, so we are leaving it alone for a while instead of adding to its load. Retry-After and `backoff.openUntil` say when; the next read after that is a single trial.
NOTHING_TO_DECODE 422 with-changes The page answered, but with no stylesheet and no text: usually a redirect stub, a parking page, or a server refusing automated reads.
NO_TOKEN_SET 422 never The site declares no usable token set, so a format built from roles cannot be produced. The DESIGN.md still describes what was read.
NOT_FOUND 404 with-changes No theme, feature or record with that id.
INTERNAL 500 later Something failed on our side. It is not the request.

TARGET_HTTP_ERROR carries the site's own status in targetStatus, and its verdict follows it: later for a 5xx, with-changes for a 4xx.

We obey robots.txt

Our reader identifies as CSSCremeBot, and it reads /robots.txt before it reads anything else on a host. It follows RFC 9309: the rules for CSSCremeBot apply when the file names us, the rules for * apply when it does not, the longest matching rule wins and Allow wins a tie. A missing file (any 4xx) means no restrictions. A file that answers with a server error means a full disallow until it answers, which is the ROBOTS_UNREACHABLE code.

It applies to every fetch our server makes from someone else's site: the page, every redirect hop, and every stylesheet, where a stylesheet on a CDN is governed by the CDN's own robots.txt. The browser workspace at /decode goes through the same check, so the frame is never a way round a refusal, and a refused site is not photographed by the screenshot fallback either. The one exception is a site owner asking us to verify a token they placed on their own site, because the owner is the one asking.

To keep us off a whole site:

User-agent: CSSCremeBot
Disallow: /

When the page is not the site

A bot challenge, a parked domain, a host's default page and a suspended account all answer with HTML and CSS of their own, and a reader that measures whatever comes back reports the registrar's design, or the vendor's, as the site's. We check first. A page that is not the site is refused with NOT_THE_SITE, which kind it is, and the evidence, so the refusal can be checked rather than taken on trust:

{
  "error": "http://brandkit.co/ is not the site: brandkit.co is showing a parked or for-sale page, not a site. Decode the address the site actually lives at.",
  "code": "NOT_THE_SITE",
  "retry": "with-changes",
  "status": 422,
  "identity": {
    "kind": "parked",
    "evidence": [
      "loads or links to a parking service (porkbun.com/account/webhosting/)",
      "the page says \"Parked at Porkbun\"",
      "title \"brandkit.co — Coming Soon\""
    ]
  }
}

The walls we recognise, from the ones measured the day this shipped:

Answered us withWallHow we know
Capterra, Glassdoor, Indeed, UpworkCloudflare challengethe cf-mitigated: challenge header
G2, EtsyDataDomea 403 with DataDome naming itself in a header or in the markup
Zillow, FiverrHUMAN (PerimeterX)a 403 carrying its press-and-hold captcha

A challenge is recognised and never solved, so its verdict is never. A vendor's script on a real page is not a wall: plenty of sites load Cloudflare's or DataDome's scripts on pages that are the site, so a script alone never decides. It takes the vendor's own header, a blocking status with its markup, or a near-empty page in a wall's own words.

When a site keeps failing

Every failure says whether to retry, and agents do. So we keep count per site: three timeouts, refused connections or 5xx answers inside two minutes, and we stop reading it for a minute and answer BACKOFF with a Retry-After. It doubles each time, up to thirty minutes, and when it expires a single read goes through as a trial before the rest. If the site sent its own Retry-After, we wait at least that long. A 404, a robots refusal or a wall never counts: those say nothing about strain. This protects the site, not us.

What every read cost the site

Every decode carries work: each request it made, how many went to the site's own host and how many to a CDN, the robots.txt answers it reused, redirects, bytes and time, with the log. Nothing in it is estimated. A real one, stripe.com:

{
  "requests": 8,
  "toSite": 2,
  "toOtherHosts": 6,
  "hosts": [
    {
      "host": "b.stripecdn.com",
      "requests": 6
    },
    {
      "host": "stripe.com",
      "requests": 2
    }
  ],
  "robots": {
    "read": 2,
    "reused": 5,
    "ms": 635
  },
  "redirects": 0,
  "stylesheets": {
    "planned": 5,
    "fetched": 5,
    "failed": 0,
    "refusedByRobots": 0,
    "skippedOverBudget": 0
  },
  "bytes": {
    "html": 656359,
    "css": 488267,
    "robots": 886,
    "total": 1145512
  },
  "ms": {
    "total": 1045,
    "page": 169,
    "stylesheets": 396
  },
  "log": [
    {
      "kind": "robots",
      "url": "stripe.com/robots.txt",
      "status": 200,
      "ms": 361,
      "bytes": 643
    },
    {
      "kind": "page",
      "url": "stripe.com/",
      "status": 200,
      "ms": 169,
      "bytes": 656359
    },
    {
      "kind": "robots",
      "url": "b.stripecdn.com/robots.txt",
      "status": 403,
      "ms": 274,
      "bytes": 243
    },
    {
      "kind": "stylesheet",
      "url": "b.stripecdn.com/mkt-ssr-statics/assets/_next/static/css/5f8c822d1f0f2d3e.css",
      "status": 200,
      "ms": 59,
      "bytes": 2723
    }
  ]
}

A read is also reused for ten minutes by default, so asking for the JSON, the DESIGN.md and tokens.json costs the site one read, and all three carry the same fingerprint. maxAge sets how old is acceptable (up to an hour; 0 reads the site now). A reused read keeps its own fetchedAt, says cache.hit: true and how old it is, and its work shows zero requests with the original receipt beside it.

Checking a brand spec

Pass expect as an object and each field is checked on its own: the colour roles, the typefaces, the radius, the base and root size, the spacing grid, dark mode, and named tokens. Every field comes back match, near, different or not-declared, with what it was compared against. not-declared means the site gives nothing to compare, and it is never filled with the nearest thing in sight. The summary counts verdicts; there is no score.

https://csscreme.com/api/decode-url?url=https://stripe.com&expect={"primary":"#635BFF","bg":"#FFFFFF","radius":"4px","spacing":"4px"}

What a successful read carries

The same honesty on the way in. Every decode now says when it was fetched, what robots.txt said, and what 1rem means on that site. The last one matters more than it sounds: a site that sets html { font-size: 62.5% } has 1rem = 10px, and a reader that assumes 16 prints every rem size 1.6 times too large. Each spacing and type-scale row keeps the value as written beside the pixels it resolved to, so 32 can be checked against the 3.2rem in the stylesheet. A list that comes back empty is named in notMeasured with the reason, rather than looking like a site that declares nothing.

{
  "fetchedAt": "2026-09-25T12:51:05.675Z",
  "robots": {
    "checked": true,
    "allowed": true,
    "rule": null,
    "source": "https://pranksy.io/robots.txt",
    "status": 404,
    "note": "No robots.txt (HTTP 404), which RFC 9309 reads as no restrictions."
  },
  "root": {
    "px": 10,
    "declared": true,
    "from": "html { font-size: 62.5% }",
    "em": {
      "px": 10,
      "from": "the root (no body font-size declared)"
    }
  },
  "spacing": [
    {
      "value": "3.2rem",
      "n": 14,
      "unit": "rem",
      "px": 32,
      "rootPx": 10
    }
  ]
}