AI design systems

What Google's DESIGN.md Linter Finds in Real Public DESIGN.md Files

48 public DESIGN.md files, one linter version, every finding logged. What the rules catch, what slips through and how to read a clean run.

· Diagram Studio Editorial

Direct answer

In a non-representative sample of 48 public files run on linter 0.4.0 (October 1, 2026), 21 exited with errors, mostly CSS values like clamp() that DESIGN.md dimensions reject. Orphaned tokens were the most common warning. The linter checks tokens and structure, not whether prose matches tokens, so a clean result is not proof of a coherent design system.

The short answer

We ran Google's DESIGN.md linter (@google/design.md 0.4.0) on 48 public DESIGN.md files from 46 GitHub repositories on October 1, 2026. Twenty-one files exited with an error. Eight finished with no errors or warnings. The most common error was not a design mistake. It was a CSS value the format does not accept as a dimension: clamp() font sizes, normal letter spacing, and a unitless 0.

The most common warning was an orphaned token, a color defined but never used by a component, which appeared in 28 files. The linter did not flag prose that contradicted the tokens, an emptied Colors section, a 2px display font, or a duplicate section heading that the spec says should be rejected.

The practical reading: the linter is a token and structure validator. A clean run tells you the tokens parsed and resolve. It does not tell you the design system is coherent, and it can occasionally tell you nothing at all, because one sampled file produced no findings and no token summary since its tokens were never read.

What a DESIGN.md file is

DESIGN.md is a plain-text format for describing a visual identity to coding agents. Google Labs open-sourced the draft specification on April 21, 2026, as announced on the Google blog and by the Stitch account. The repository is Apache-2.0 licensed and the format is labelled alpha.

Today, we’re open-sourcing the draft specification for DESIGN.md, so it can be used across any tool or platform. We’re also adding new capabilities. DESIGN.md lets you easily export and import your design rules from project to project. Instead of guessing intent, agents know

— Stitch by Google (@stitchbygoogle) View on X
The announcement that put the draft spec in public, dated the same day as the first npm release of the CLI.

A file has two parts. The first is optional YAML front matter holding design tokens: named values for colors, typography, rounded, spacing and components. Tokens can reference each other with curly-brace paths such as {colors.primary}. The second part is Markdown prose in a fixed set of ## sections (Overview, Colors, Typography, Layout, Elevation & Depth, Shapes, Components, Do's and Don'ts) that explains why the values exist.

The tokens are the normative values; the prose provides context for how to apply them.

DESIGN.md specification, google-labs-code/design.md, docs/spec.md — DESIGN.md specification (docs/spec.md)

That sentence matters for what follows. The tokens are the part a tool can verify, so a linter has a clear job there. The prose is the part that carries intent, and nothing in the spec requires it to agree with the tokens.

Watch Meet DESIGN.md: A new open standard for AI-generated UI by Google for Developers on YouTube
A video introduction to the format from Google for Developers, useful if you want the format explained before reading the linter results.

What the linter is documented to check

The CLI has four commands: lint, diff, export and spec. design.md lint <file> prints a JSON report of findings, each with a severity, and exits with code 1 when there is at least one error, otherwise 0. In 0.4.0, passing --format text still printed JSON in our runs.

The repository README lists eleven rules. The table adds what happened in our sample of 48 files. Severities in the second column are the documented ones.

Documented lint rules and what each did across 48 real files (v0.4.0, run October 1, 2026). Counts are findings / files affected.
RuleDocumented severityWhat it checks (documented)In the 48 real files (logged)
broken-referrorReferences that do not resolve38 / 3 files. 2 errors (unresolved reference, 1 file); 36 warnings (unrecognized component sub-token, 3 files)
missing-primarywarningColors defined, no primary24 files
contrast-ratiowarningComponent text/background pairs under 4.5:137 / 13 files
orphaned-tokenswarningColors never referenced by a component303 / 28 files
token-summaryinfoCounts of tokens per section45 files
missing-sectionsinfoNo spacing or rounded section11 / 9 files
missing-typographywarningColors but no typographyNever fired (fired in our probe)
section-orderwarningSections out of canonical orderNever fired (fired in our probes)
unknown-keywarningTop-level key that looks like a typoNever fired (fired in our probe)
token-like-ignoredwarningUnknown top-level key holding token-like values10 / 4 files
omitted-rulesinfoChecks the omitted configurationNever fired (not probed)

Two details in that table are discrepancies between the documentation and the behaviour we saw. First, broken-ref is documented as an error, but 36 of its 38 findings were warnings: it also covers component properties the spec does not list (fontSize, borderRadius, minHeight and similar). Only the 2 true unresolved references were errors. Second, the installed 0.4.0 package README says "seven rules" yet lists eight, while the main-branch README says eleven; the package we ran does contain the extra rule IDs. The README appears to be a version behind the code.

The table also leaves out a class of findings. 182 findings in our runs had no rule ID at all: they come from the parser, not from a named rule. They include invalid dimensions, invalid colors, unrecognized typography properties and YAML errors, and they carry 96 of the 98 errors we recorded. If you filter CI output by rule ID, you would miss most of the errors that actually failed our files.

How the 48 files were chosen and run

Selection was mechanical, which makes it reproducible and also biased. We searched GitHub code for files named DESIGN.md containing both colors: and typography:, took the results in the search's default order, kept the first file from each distinct repository, and stopped at 45 third-party repositories. We added the three example files in Google's own repository. That makes 48 files in 46 repositories.

The query requires token-like text, so files with no tokens never entered the sample. A few files still turned out to contain none: one is a lowercase design.md about a compression library rather than a design system. We kept it and report what the linter said.

Each file was downloaded at the commit SHA shown in the search hit, linted in a scratch directory outside any project, and written out as JSON. The manifest with repository, path, SHA and file hashes, the raw JSON output per file, the scripts, and the Node (22.23.1) and linter (0.4.0) versions are in the article's evidence folder. We do not know who or what wrote any of these files, and we make no claim about it.

  • **Command:** npx design.md lint --format json <file> from a directory where @google/design.md@0.4.0 was installed.
  • **Package:** version 0.4.0 was the latest on npm and was published July 27, 2026; the first release, 0.1.0, was published April 21, 2026.
  • **Result:** 21 of 48 exited 1; 27 exited 0. Across all files: 98 errors and 496 warnings.

What the linter flagged most

Hand-drawn bar chart of five finding kinds by number of files affected out of 48: orphaned tokens 28, missing primary 24, invalid dimension 20, contrast ratio 13, missing sections 9.
Files affected by each kind of finding in the 48-file sample; invalid dimension comes from the parser rather than a named rule.

**Invalid dimensions failed 20 of the 21 failing files.** The spec limits a dimension to a number with px, em or rem. Many files use values that are valid CSS but not valid here. By file: clamp() font sizes in 14 files, normal letter spacing in 9, a bare 0 in 4, and calc() or var() expressions in 3. One file also wrote 24 of its color values as var(--color-name) references, which produced 24 errors. The one remaining failing file had two unresolved token references.

This points to a gap between how designers write type scales and what the format can store. A fluid clamp() size cannot be expressed as a single dimension token. The linter is right to reject it under the spec, but the files suggest authors reach for CSS because they are describing a CSS design system. We tested a mechanical fix on three failing files: replacing normal and 0 with 0em and each clamp() with its maximum value cleared the errors in two of the three. Whether the maximum is the right value to hand an agent is a design decision, not a lint one.

**Orphaned tokens appeared in 28 of the 45 files whose tokens parsed.** The rule only compares colors with component references. Files with a full palette and a handful of components trip it constantly: one file defined 42 colors against 11 components and received 31 warnings. For a palette meant to be used by agents directly, an unreferenced color may be deliberate, so treat this rule as a map of which colors your components do not mention, not as a defect count.

**Missing primary appeared in 24 files.** The rule looks for a color literally named primary. Several of those files do define a brand color under another name, such as brand-primary, primary-ink or brand-green, and the linter still warns that agents will invent one. That advice follows the spec's convention, but the finding does not distinguish a file with no brand color from one that simply calls it something else.

**Contrast findings came in 13 files, and most are suspect.** The rule flags component pairs under the WCAG AA minimum of 4.5:1 for normal text (WCAG 2.2, 1.4.3). Of 37 contrast findings, 27 in 9 files involved a background color with an alpha channel, such as #ffffff1a or #00000000. In a direct probe, black text on #ffffff00 (fully transparent) produced no warning, which indicates the check ignores alpha and compares only the RGB values. The WCAG page says nothing about alpha, so what a "correct" result would be depends on what sits behind the element, which a token file cannot say. Treat contrast results on translucent colors as unreliable.

Smaller groups: token-like-ignored in 4 files (for example top-level motion, strokes or signals maps that exports silently drop), and 84 warnings in 5 files for typography properties the schema does not know, such as textTransform and fontStyle.

What the linter did not flag

Counting what is absent needs controlled input, so we copied Google's paws-and-paths example, which linted with no errors or warnings, and changed one thing at a time. These are our own probes of one example on 0.4.0; they show what the linter does with these inputs, not everything it can catch.

Two-column hand-drawn comparison: five deliberate edits the linter stayed silent about, and four edits it flagged, including a failing contrast pair and an unresolved reference.
Results of single-change probes on one clean Google example; silent means no finding beyond the token summary.
Single-change probes on Google's paws-and-paths example, linter 0.4.0, October 1, 2026.
ChangeLinter result
Prose says primary is bright red #FF0000; token unchangedNo finding
All prose after the front matter deletedNo finding
Colors section emptiedNo finding
Display font size set to 2pxNo finding
fontWeight: 12345No finding
Prose tells readers to use {colors.brand-magenta}, which does not existNo finding
Second ## Colors section appendedOne section-order warning, exit 0; the spec says duplicates are an error
Shapes section moved before Colorssection-order warning
Failing component color pair addedcontrast-ratio warning
full: 50% for a rounded valueError with no rule ID (invalid unit)
typography: misspelled as typograhpy:4 errors, 3 warnings: unresolved references, missing-typography, unknown-key and token-like-ignored
Reference to {colors.nonexistent}broken-ref error

The pattern is consistent. The linter checks that tokens parse, resolve and meet one accessibility threshold. It does not compare prose with tokens, judge whether a value is sensible, or (as the duplicate-heading probe shows) enforce every rule the spec states. Spec-versus-linter differences are expected in an alpha project, but a team that assumes the linter implements the spec in full will be wrong in at least this one case.

The misspelled key shows the opposite effect. One typo produced seven findings, because every component reference to a typography token then failed. When a lint run is noisy, look for a single upstream cause first.

**Silent passes.** Real files showed the sharpest case. The TrekSphere file begins with a different front-matter block (trigger: always_on) followed by a second block whose indentation has been flattened. The linter returned an empty findings list: no errors, no warnings and, notably, no token summary. A second file, the compression-library design.md, returned only a "No YAML content found" warning. A third had a YAML syntax error and also returned no summary. In all three the linter read no tokens. All three exited 0. An independent report from September 2026 found the same "No YAML content found" result on a sample file from a third-party site when run with the same 0.4.0 release (reported).

Across the 48 real files, section-order, missing-typography, unknown-key and omitted-rules never fired. We cannot say whether that reflects careful authors, a mostly template-shaped sample, or rules that rarely apply; the probes show at least three of the four do work when triggered.

How to read a lint run

The sample suggests a short triage order. It is our inference from these runs, not an official procedure.

Hand-drawn flowchart for reading a lint run: check exit code, fix errors, confirm a token summary exists, check counts match the file, triage warnings, review prose by hand.
A triage order drawn from the sample: errors, then proof that tokens were read, then warnings, then a human read of the prose.
  1. Pin the version, for example npx @google/design.md@0.4.0 lint --format json DESIGN.md, and save the JSON. The rules and severities are changing across releases.
  2. Fix errors first, and expect most to be value formats: add units, replace normal and clamp(), and resolve references.
  3. Confirm a token-summary finding exists and its counts match what you wrote. Three of our 48 files had none, and in those the linter had read no tokens.
  4. Read warnings by rule. Treat contrast-ratio on translucent colors with suspicion and orphaned-tokens as information.
  5. Review the prose against the tokens by hand, since no rule compares them.

Worked example: a file defines colors with primary: "#1A1C1E" but the Colors prose says the primary is a warm red. The linter returns the token summary and nothing else. Only step 5 catches it. The agent reading the file receives both statements, and the spec says the tokens are normative, so the red in the prose is likely to be ignored or to confuse the result. That outcome is inference; we did not test how any agent reacts.

What this sample cannot tell you

It cannot tell you how common these problems are in all DESIGN.md files. The selection came from one search, one ranking and one date. Newer or differently named files, private repositories and files without tokens are outside it.

It cannot tell you whether an agent produces better interfaces from a file that lints clean. We did not test agent output, and the sources we found did not measure it either, so claims that DESIGN.md improves generated UI are unverified here.

It cannot tell you who wrote the files. Many contain generic names and similar structures, but similar structure is not evidence of origin.

It also cannot predict future versions. 0.4.0 was released on July 27, 2026, and five versions (0.1.0 to 0.4.0) were published between April and July 2026. Re-run the commands in the evidence folder against the current release before relying on any count above.

Is the linter worth running?

Yes, for what it does. It catches unresolved references, malformed values and unreadable front matter before an agent or exporter trips over them, and in our sample 21 of 48 files had at least one such error. That is a useful gate on the machine-readable half of the file.

It is not a review of the design system. Prose that contradicts tokens, implausible values and duplicate headings can all pass, and an unreadable file can pass with no findings. Run it in CI with a pinned version, require a token summary, and keep a human check on the prose. Adopting DESIGN.md is cheap to try; trusting a clean lint result as proof of coherence is the mistake the sample warns against.