AI design systems
What Google's DESIGN.md Linter Finds in Real Public DESIGN.md Files
48 public DESIGN.md files, one linter version, every finding logged. What the rules catch, what slips through and how to read a clean run.
· Diagram Studio Editorial
Direct answer
In a non-representative sample of 48 public files run on linter 0.4.0 (October 1, 2026), 21 exited with errors, mostly CSS values like clamp() that DESIGN.md dimensions reject. Orphaned tokens were the most common warning. The linter checks tokens and structure, not whether prose matches tokens, so a clean result is not proof of a coherent design system.
The short answer
We ran Google's DESIGN.md linter (@google/design.md 0.4.0) on 48 public DESIGN.md files from 46 GitHub repositories on October 1, 2026. Twenty-one files exited with an error. Eight finished with no errors or warnings. The most common error was not a design mistake. It was a CSS value the format does not accept as a dimension: clamp() font sizes, normal letter spacing, and a unitless 0.
The most common warning was an orphaned token, a color defined but never used by a component, which appeared in 28 files. The linter did not flag prose that contradicted the tokens, an emptied Colors section, a 2px display font, or a duplicate section heading that the spec says should be rejected.
The practical reading: the linter is a token and structure validator. A clean run tells you the tokens parsed and resolve. It does not tell you the design system is coherent, and it can occasionally tell you nothing at all, because one sampled file produced no findings and no token summary since its tokens were never read.
What a DESIGN.md file is
DESIGN.md is a plain-text format for describing a visual identity to coding agents. Google Labs open-sourced the draft specification on April 21, 2026, as announced on the Google blog and by the Stitch account. The repository is Apache-2.0 licensed and the format is labelled alpha.
Today, we’re open-sourcing the draft specification for DESIGN.md, so it can be used across any tool or platform. We’re also adding new capabilities. DESIGN.md lets you easily export and import your design rules from project to project. Instead of guessing intent, agents know
— Stitch by Google (@stitchbygoogle) View on X
A file has two parts. The first is optional YAML front matter holding design tokens: named values for colors, typography, rounded, spacing and components. Tokens can reference each other with curly-brace paths such as {colors.primary}. The second part is Markdown prose in a fixed set of ## sections (Overview, Colors, Typography, Layout, Elevation & Depth, Shapes, Components, Do's and Don'ts) that explains why the values exist.
The tokens are the normative values; the prose provides context for how to apply them.
That sentence matters for what follows. The tokens are the part a tool can verify, so a linter has a clear job there. The prose is the part that carries intent, and nothing in the spec requires it to agree with the tokens.
What the linter is documented to check
The CLI has four commands: lint, diff, export and spec. design.md lint <file> prints a JSON report of findings, each with a severity, and exits with code 1 when there is at least one error, otherwise 0. In 0.4.0, passing --format text still printed JSON in our runs.
The repository README lists eleven rules. The table adds what happened in our sample of 48 files. Severities in the second column are the documented ones.
| Rule | Documented severity | What it checks (documented) | In the 48 real files (logged) |
|---|---|---|---|
| broken-ref | error | References that do not resolve | 38 / 3 files. 2 errors (unresolved reference, 1 file); 36 warnings (unrecognized component sub-token, 3 files) |
| missing-primary | warning | Colors defined, no primary | 24 files |
| contrast-ratio | warning | Component text/background pairs under 4.5:1 | 37 / 13 files |
| orphaned-tokens | warning | Colors never referenced by a component | 303 / 28 files |
| token-summary | info | Counts of tokens per section | 45 files |
| missing-sections | info | No spacing or rounded section | 11 / 9 files |
| missing-typography | warning | Colors but no typography | Never fired (fired in our probe) |
| section-order | warning | Sections out of canonical order | Never fired (fired in our probes) |
| unknown-key | warning | Top-level key that looks like a typo | Never fired (fired in our probe) |
| token-like-ignored | warning | Unknown top-level key holding token-like values | 10 / 4 files |
| omitted-rules | info | Checks the omitted configuration | Never fired (not probed) |
Two details in that table are discrepancies between the documentation and the behaviour we saw. First, broken-ref is documented as an error, but 36 of its 38 findings were warnings: it also covers component properties the spec does not list (fontSize, borderRadius, minHeight and similar). Only the 2 true unresolved references were errors. Second, the installed 0.4.0 package README says "seven rules" yet lists eight, while the main-branch README says eleven; the package we ran does contain the extra rule IDs. The README appears to be a version behind the code.
The table also leaves out a class of findings. 182 findings in our runs had no rule ID at all: they come from the parser, not from a named rule. They include invalid dimensions, invalid colors, unrecognized typography properties and YAML errors, and they carry 96 of the 98 errors we recorded. If you filter CI output by rule ID, you would miss most of the errors that actually failed our files.
How the 48 files were chosen and run
Selection was mechanical, which makes it reproducible and also biased. We searched GitHub code for files named DESIGN.md containing both colors: and typography:, took the results in the search's default order, kept the first file from each distinct repository, and stopped at 45 third-party repositories. We added the three example files in Google's own repository. That makes 48 files in 46 repositories.
The query requires token-like text, so files with no tokens never entered the sample. A few files still turned out to contain none: one is a lowercase design.md about a compression library rather than a design system. We kept it and report what the linter said.
Each file was downloaded at the commit SHA shown in the search hit, linted in a scratch directory outside any project, and written out as JSON. The manifest with repository, path, SHA and file hashes, the raw JSON output per file, the scripts, and the Node (22.23.1) and linter (0.4.0) versions are in the article's evidence folder. We do not know who or what wrote any of these files, and we make no claim about it.
- **Command:**
npx design.md lint --format json <file>from a directory where@google/design.md@0.4.0was installed. - **Package:** version 0.4.0 was the latest on npm and was published July 27, 2026; the first release, 0.1.0, was published April 21, 2026.
- **Result:** 21 of 48 exited 1; 27 exited 0. Across all files: 98 errors and 496 warnings.
What the linter flagged most

**Invalid dimensions failed 20 of the 21 failing files.** The spec limits a dimension to a number with px, em or rem. Many files use values that are valid CSS but not valid here. By file: clamp() font sizes in 14 files, normal letter spacing in 9, a bare 0 in 4, and calc() or var() expressions in 3. One file also wrote 24 of its color values as var(--color-name) references, which produced 24 errors. The one remaining failing file had two unresolved token references.
This points to a gap between how designers write type scales and what the format can store. A fluid clamp() size cannot be expressed as a single dimension token. The linter is right to reject it under the spec, but the files suggest authors reach for CSS because they are describing a CSS design system. We tested a mechanical fix on three failing files: replacing normal and 0 with 0em and each clamp() with its maximum value cleared the errors in two of the three. Whether the maximum is the right value to hand an agent is a design decision, not a lint one.
**Orphaned tokens appeared in 28 of the 45 files whose tokens parsed.** The rule only compares colors with component references. Files with a full palette and a handful of components trip it constantly: one file defined 42 colors against 11 components and received 31 warnings. For a palette meant to be used by agents directly, an unreferenced color may be deliberate, so treat this rule as a map of which colors your components do not mention, not as a defect count.
**Missing primary appeared in 24 files.** The rule looks for a color literally named primary. Several of those files do define a brand color under another name, such as brand-primary, primary-ink or brand-green, and the linter still warns that agents will invent one. That advice follows the spec's convention, but the finding does not distinguish a file with no brand color from one that simply calls it something else.
**Contrast findings came in 13 files, and most are suspect.** The rule flags component pairs under the WCAG AA minimum of 4.5:1 for normal text (WCAG 2.2, 1.4.3). Of 37 contrast findings, 27 in 9 files involved a background color with an alpha channel, such as #ffffff1a or #00000000. In a direct probe, black text on #ffffff00 (fully transparent) produced no warning, which indicates the check ignores alpha and compares only the RGB values. The WCAG page says nothing about alpha, so what a "correct" result would be depends on what sits behind the element, which a token file cannot say. Treat contrast results on translucent colors as unreliable.
Smaller groups: token-like-ignored in 4 files (for example top-level motion, strokes or signals maps that exports silently drop), and 84 warnings in 5 files for typography properties the schema does not know, such as textTransform and fontStyle.
What the linter did not flag
Counting what is absent needs controlled input, so we copied Google's paws-and-paths example, which linted with no errors or warnings, and changed one thing at a time. These are our own probes of one example on 0.4.0; they show what the linter does with these inputs, not everything it can catch.

| Change | Linter result |
|---|---|
Prose says primary is bright red #FF0000; token unchanged | No finding |
| All prose after the front matter deleted | No finding |
| Colors section emptied | No finding |
Display font size set to 2px | No finding |
fontWeight: 12345 | No finding |
Prose tells readers to use {colors.brand-magenta}, which does not exist | No finding |
Second ## Colors section appended | One section-order warning, exit 0; the spec says duplicates are an error |
| Shapes section moved before Colors | section-order warning |
| Failing component color pair added | contrast-ratio warning |
full: 50% for a rounded value | Error with no rule ID (invalid unit) |
typography: misspelled as typograhpy: | 4 errors, 3 warnings: unresolved references, missing-typography, unknown-key and token-like-ignored |
Reference to {colors.nonexistent} | broken-ref error |
The pattern is consistent. The linter checks that tokens parse, resolve and meet one accessibility threshold. It does not compare prose with tokens, judge whether a value is sensible, or (as the duplicate-heading probe shows) enforce every rule the spec states. Spec-versus-linter differences are expected in an alpha project, but a team that assumes the linter implements the spec in full will be wrong in at least this one case.
The misspelled key shows the opposite effect. One typo produced seven findings, because every component reference to a typography token then failed. When a lint run is noisy, look for a single upstream cause first.
**Silent passes.** Real files showed the sharpest case. The TrekSphere file begins with a different front-matter block (trigger: always_on) followed by a second block whose indentation has been flattened. The linter returned an empty findings list: no errors, no warnings and, notably, no token summary. A second file, the compression-library design.md, returned only a "No YAML content found" warning. A third had a YAML syntax error and also returned no summary. In all three the linter read no tokens. All three exited 0. An independent report from September 2026 found the same "No YAML content found" result on a sample file from a third-party site when run with the same 0.4.0 release (reported).
Across the 48 real files, section-order, missing-typography, unknown-key and omitted-rules never fired. We cannot say whether that reflects careful authors, a mostly template-shaped sample, or rules that rarely apply; the probes show at least three of the four do work when triggered.
How to read a lint run
The sample suggests a short triage order. It is our inference from these runs, not an official procedure.

- Pin the version, for example
npx @google/design.md@0.4.0 lint --format json DESIGN.md, and save the JSON. The rules and severities are changing across releases. - Fix errors first, and expect most to be value formats: add units, replace
normalandclamp(), and resolve references. - Confirm a
token-summaryfinding exists and its counts match what you wrote. Three of our 48 files had none, and in those the linter had read no tokens. - Read warnings by rule. Treat
contrast-ratioon translucent colors with suspicion andorphaned-tokensas information. - Review the prose against the tokens by hand, since no rule compares them.
Worked example: a file defines colors with primary: "#1A1C1E" but the Colors prose says the primary is a warm red. The linter returns the token summary and nothing else. Only step 5 catches it. The agent reading the file receives both statements, and the spec says the tokens are normative, so the red in the prose is likely to be ignored or to confuse the result. That outcome is inference; we did not test how any agent reacts.
What this sample cannot tell you
It cannot tell you how common these problems are in all DESIGN.md files. The selection came from one search, one ranking and one date. Newer or differently named files, private repositories and files without tokens are outside it.
It cannot tell you whether an agent produces better interfaces from a file that lints clean. We did not test agent output, and the sources we found did not measure it either, so claims that DESIGN.md improves generated UI are unverified here.
It cannot tell you who wrote the files. Many contain generic names and similar structures, but similar structure is not evidence of origin.
It also cannot predict future versions. 0.4.0 was released on July 27, 2026, and five versions (0.1.0 to 0.4.0) were published between April and July 2026. Re-run the commands in the evidence folder against the current release before relying on any count above.
Is the linter worth running?
Yes, for what it does. It catches unresolved references, malformed values and unreadable front matter before an agent or exporter trips over them, and in our sample 21 of 48 files had at least one such error. That is a useful gate on the machine-readable half of the file.
It is not a review of the design system. Prose that contradicts tokens, implausible values and duplicate headings can all pass, and an unreadable file can pass with no findings. Run it in CI with a pinned version, require a token summary, and keep a human check on the prose. Adopting DESIGN.md is cheap to try; trusting a clean lint result as proof of coherence is the mistake the sample warns against.