For agents calling the decoder
Every failure says whether to try again.
A sentence tells a person what went wrong. An agent needs one more thing: should it retry,
change the request, or stop? Every error from /api/decode-url, the
MCP tools, the legibility check and the extractor carries
a code from the list below and a retry verdict beside the sentence. Existing
readers of error keep working: it is still the sentence.
What a refusal looks like
A real one. Google's robots.txt disallows /search for every crawler, so we do not read it:
{
"error": "Could not read https://www.google.com/search?q=design: robots.txt disallows CSSCremeBot on /search (Disallow: /search)",
"code": "ROBOTS_DENIED",
"retry": "never",
"status": 403,
"hint": "The site owner asked crawlers not to read this path, and we obey that everywhere, including the /decode workspace. Nothing on our side will change the answer.",
"robots": {
"checked": true,
"allowed": false,
"agent": "*",
"rule": "Disallow: /search",
"source": "https://www.google.com/robots.txt",
"status": 200
},
"errors": "https://csscreme.com/ai/errors"
} Over MCP the same failure is a tool error whose text leads with the code and the verdict, so the first
token is enough to branch on. The pair is also in _meta["com.csscreme/error"] as data.
[ROBOTS_DENIED] retry: never. Could not read https://www.google.com/search?q=design: robots.txt disallows CSSCremeBot on /search (Disallow: /search). … The three retry words
never- The same request will fail the same way, and so will any rewording of it. Stop.
later- The same request may succeed after a wait. When we know how long, Retry-After says so.
with-changes- This request will keep failing. A corrected one (another URL, a valid parameter) may not.
Every code
| Code | HTTP | Retry | What it means |
|---|---|---|---|
MISSING_URL | 400 | with-changes | No url was given. |
INVALID_URL | 400 | with-changes | Not an absolute http(s) URL on a standard port. |
BLOCKED_ADDRESS | 400 | never | The host is, or redirects to, a private, loopback, link-local or cloud-metadata address. Refused before anything is fetched. |
INVALID_PARAM | 400 | with-changes | A parameter other than the URL is malformed or unknown: a format name, an expect colour, a page count. |
METHOD_NOT_ALLOWED | 405 | with-changes | The endpoint answers a different HTTP method. |
RATE_LIMITED | 429 | later | Too many requests from one address in a minute. Retry-After says how long to wait. |
ROBOTS_DENIED | 403 | never | The site's robots.txt disallows CSSCremeBot on this path, and we obey it. Retrying will not change that, and neither will another CSS Crème tool: they all read through the same check. |
ROBOTS_UNREACHABLE | 503 | later | The site's robots.txt answered with a server error or could not be reached. RFC 9309 says to treat that as a full disallow until it answers. |
TARGET_HTTP_ERROR | 502 | with-changes | The site answered with an HTTP error status. The target status is in `targetStatus`; a 5xx there may clear later, a 4xx will not. |
TIMEOUT | 504 | later | The page or its stylesheets did not arrive within the time budget, or the host did not accept a connection in time. |
HOST_NOT_FOUND | 400 | with-changes | The host name does not resolve in DNS: a typo, a domain that does not exist, or one that has lapsed. Nothing was fetched. |
TLS_FAILED | 502 | with-changes | The site's HTTPS handshake failed. Many parked and old domains only answer plain http; if that is the address, pass the http:// URL. |
FETCH_FAILED | 502 | later | The connection failed before an answer arrived: DNS, TLS, a reset or a refused connection. |
NOT_THE_SITE | 422 | never | The page that answered is not the site: a bot challenge, a parked or for-sale domain, a host's default page, or a suspended account. `identity` names which, and the evidence. A challenge is never solved, so its verdict is never; a parked domain's is with-changes (decode the address the site really lives at). |
BACKOFF | 503 | later | This site failed several reads in a row, so we are leaving it alone for a while instead of adding to its load. Retry-After and `backoff.openUntil` say when; the next read after that is a single trial. |
NOTHING_TO_DECODE | 422 | with-changes | The page answered, but with no stylesheet and no text: usually a redirect stub, a parking page, or a server refusing automated reads. |
NO_TOKEN_SET | 422 | never | The site declares no usable token set, so a format built from roles cannot be produced. The DESIGN.md still describes what was read. |
NOT_FOUND | 404 | with-changes | No theme, feature or record with that id. |
INTERNAL | 500 | later | Something failed on our side. It is not the request. |
TARGET_HTTP_ERROR carries the site's own status in targetStatus,
and its verdict follows it: later for a 5xx, with-changes for a 4xx.
We obey robots.txt
Our reader identifies as CSSCremeBot, and it reads /robots.txt before it reads
anything else on a host. It follows RFC 9309:
the rules for CSSCremeBot apply when the file names us, the rules for * apply when it
does not, the longest matching rule wins and Allow wins a tie. A missing file (any 4xx) means no
restrictions. A file that answers with a server error means a full disallow until it answers, which is the
ROBOTS_UNREACHABLE code.
It applies to every fetch our server makes from someone else's site: the page, every redirect hop, and every stylesheet, where a stylesheet on a CDN is governed by the CDN's own robots.txt. The browser workspace at /decode goes through the same check, so the frame is never a way round a refusal, and a refused site is not photographed by the screenshot fallback either. The one exception is a site owner asking us to verify a token they placed on their own site, because the owner is the one asking.
To keep us off a whole site:
User-agent: CSSCremeBot
Disallow: / When the page is not the site
A bot challenge, a parked domain, a host's default page and a suspended account all answer with HTML and CSS
of their own, and a reader that measures whatever comes back reports the registrar's design, or the vendor's,
as the site's. We check first. A page that is not the site is refused with NOT_THE_SITE, which kind
it is, and the evidence, so the refusal can be checked rather than taken on trust:
{
"error": "http://brandkit.co/ is not the site: brandkit.co is showing a parked or for-sale page, not a site. Decode the address the site actually lives at.",
"code": "NOT_THE_SITE",
"retry": "with-changes",
"status": 422,
"identity": {
"kind": "parked",
"evidence": [
"loads or links to a parking service (porkbun.com/account/webhosting/)",
"the page says \"Parked at Porkbun\"",
"title \"brandkit.co — Coming Soon\""
]
}
} The walls we recognise, from the ones measured the day this shipped:
| Answered us with | Wall | How we know |
|---|---|---|
| Capterra, Glassdoor, Indeed, Upwork | Cloudflare challenge | the cf-mitigated: challenge header |
| G2, Etsy | DataDome | a 403 with DataDome naming itself in a header or in the markup |
| Zillow, Fiverr | HUMAN (PerimeterX) | a 403 carrying its press-and-hold captcha |
A challenge is recognised and never solved, so its verdict is never. A vendor's script on a real
page is not a wall: plenty of sites load Cloudflare's or DataDome's scripts on pages that are the site, so a script
alone never decides. It takes the vendor's own header, a blocking status with its markup, or a near-empty page in a
wall's own words.
When a site keeps failing
Every failure says whether to retry, and agents do. So we keep count per site: three timeouts, refused
connections or 5xx answers inside two minutes, and we stop reading it for a minute and answer BACKOFF
with a Retry-After. It doubles each time, up to thirty minutes, and when it expires a single read goes
through as a trial before the rest. If the site sent its own Retry-After, we wait at least that long. A 404, a robots
refusal or a wall never counts: those say nothing about strain. This protects the site, not us.
What every read cost the site
Every decode carries work: each request it made, how many went to the site's own host and how many
to a CDN, the robots.txt answers it reused, redirects, bytes and time, with the log. Nothing in it is estimated.
A real one, stripe.com:
{
"requests": 8,
"toSite": 2,
"toOtherHosts": 6,
"hosts": [
{
"host": "b.stripecdn.com",
"requests": 6
},
{
"host": "stripe.com",
"requests": 2
}
],
"robots": {
"read": 2,
"reused": 5,
"ms": 635
},
"redirects": 0,
"stylesheets": {
"planned": 5,
"fetched": 5,
"failed": 0,
"refusedByRobots": 0,
"skippedOverBudget": 0
},
"bytes": {
"html": 656359,
"css": 488267,
"robots": 886,
"total": 1145512
},
"ms": {
"total": 1045,
"page": 169,
"stylesheets": 396
},
"log": [
{
"kind": "robots",
"url": "stripe.com/robots.txt",
"status": 200,
"ms": 361,
"bytes": 643
},
{
"kind": "page",
"url": "stripe.com/",
"status": 200,
"ms": 169,
"bytes": 656359
},
{
"kind": "robots",
"url": "b.stripecdn.com/robots.txt",
"status": 403,
"ms": 274,
"bytes": 243
},
{
"kind": "stylesheet",
"url": "b.stripecdn.com/mkt-ssr-statics/assets/_next/static/css/5f8c822d1f0f2d3e.css",
"status": 200,
"ms": 59,
"bytes": 2723
}
]
} A read is also reused for ten minutes by default, so asking for the JSON, the DESIGN.md and tokens.json costs the
site one read, and all three carry the same fingerprint. maxAge sets how old is acceptable (up to an
hour; 0 reads the site now). A reused read keeps its own fetchedAt, says
cache.hit: true and how old it is, and its work shows zero requests with the original
receipt beside it.
Checking a brand spec
Pass expect as an object and each field is checked on its own: the colour roles, the typefaces,
the radius, the base and root size, the spacing grid, dark mode, and named tokens. Every field comes back
match, near, different or not-declared, with what it was compared
against. not-declared means the site gives nothing to compare, and it is never filled with the nearest
thing in sight. The summary counts verdicts; there is no score.
https://csscreme.com/api/decode-url?url=https://stripe.com&expect={"primary":"#635BFF","bg":"#FFFFFF","radius":"4px","spacing":"4px"} What a successful read carries
The same honesty on the way in. Every decode now says when it was fetched, what robots.txt said, and what
1rem means on that site. The last one matters more than it sounds: a site that sets
html { font-size: 62.5% } has 1rem = 10px, and a reader that assumes 16 prints every rem
size 1.6 times too large. Each spacing and type-scale row keeps the value as written beside the pixels it
resolved to, so 32 can be checked against the 3.2rem in the stylesheet. A list that
comes back empty is named in notMeasured with the reason, rather than looking like a site that
declares nothing.
{
"fetchedAt": "2026-09-25T12:51:05.675Z",
"robots": {
"checked": true,
"allowed": true,
"rule": null,
"source": "https://pranksy.io/robots.txt",
"status": 404,
"note": "No robots.txt (HTTP 404), which RFC 9309 reads as no restrictions."
},
"root": {
"px": 10,
"declared": true,
"from": "html { font-size: 62.5% }",
"em": {
"px": 10,
"from": "the root (no body font-size declared)"
}
},
"spacing": [
{
"value": "3.2rem",
"n": 14,
"unit": "rem",
"px": 32,
"rootPx": 10
}
]
}