Design systems for agents

Markdown or JSON for Design Tokens: What the Evidence Says About Which Format AI Agents Follow Better

Vendors say Markdown wins, tooling teams say JSON wins, and the controlled data is missing. Here is what each format encodes, what the evidence is worth, and how to test it on your own repository.

· Diagram Studio Editorial

Direct answer

No rigorous, reproducible comparison shows that agents follow Markdown design systems better than JSON tokens, or the reverse. JSON in the Design Tokens Community Group (DTCG) format, a W3C Community Group report, is the right source of values for build pipelines. Markdown such as DESIGN.md carries rules and rationale that raw tokens do not. Keep values in JSON, put rules in prose, and measure adherence yourself.

The short answer: they do different jobs, and nobody has measured the gap

If your question is “which file should I hand my coding agent so it stops drifting from the design system,” the honest answer is that no controlled, reproducible study published so far settles it. What exists is vendor posts, individual write-ups and research on neighbouring problems. This article lists each, says what it measured, and labels it.

What can be said with confidence is about purpose. JSON design tokens in the W3C Design Tokens format are a data interchange format: names, values and types that tools such as Style Dictionary compile into CSS, iOS and Android output. A DESIGN.md file is context for an agent: the same kind of values, plus prose on where and why to use them. Treating one as a replacement for the other is where most of the confusion starts.

A note on labels. In this article, documented means stated in an official spec or vendor page, reported means a named third party says they observed it, and inference is the editorial team's own reasoning. Nothing here comes from the editorial team running agents against these formats.

What each format is for

A design token is a named design decision stored as data, such as color.action.primary with a color value. The Design Tokens Community Group (DTCG) published its first stable Format Module, version 2025.10, on October 28, 2025. In that format every token has a $value, an optional $type, tokens sit in nested groups, and one token can alias another with {group.token} syntax. Files use the .tokens or .tokens.json extension and the media type application/design-tokens+json.

This specification is considered stable. Further updates will be provided in superseding specifications.

Design Tokens Community Group, W3C Community Group, Final Community Group Report — Design Tokens Format Module 2025.10

Stable here has a narrow meaning. The report is not a W3C Standard; it is published under the W3C Community Final Specification Agreement. Tooling is also still catching up: as of October 1, 2026, Style Dictionary's documentation says it has first-class DTCG support since version 4 but that 2025.10 is not yet fully supported and that work is in progress for version 5. Check both before you depend on them.

DESIGN.md is Google Labs' open specification for describing a visual identity to coding agents. A file has optional YAML front matter holding tokens (colors, typography, rounded corners, spacing, components) followed by Markdown sections in a fixed order: Overview, Colors, Typography, Layout, Elevation and Depth, Shapes, Components, and Do's and Don'ts. Google announced the open-sourcing on April 21, 2026. The repository states the format is at version alpha and under active development, so expect changes.

Today, we’re open-sourcing the draft specification for DESIGN.md, so it can be used across any tool or platform. We’re also adding new capabilities. DESIGN.md lets you easily export and import your design rules from project to project. Instead of guessing intent, agents know

— Stitch by Google (@stitchbygoogle) View on X
The official announcement describes DESIGN.md as a draft specification, which matches the alpha label in the repository.

The two formats are not exclusive. DESIGN.md's own spec settles which one wins inside the file, and its command-line tool can export the front matter to the W3C token format as well as to Tailwind.

The tokens are the normative values; the prose provides context for how to apply them.

Google Labs, DESIGN.md specification, google-labs-code repository — DESIGN.md specification (docs/spec.md)
Watch Meet DESIGN.md: A new open standard for AI-generated UI by Google for Developers on YouTube
Google embeds this overview in its own announcement post; watch it for the intended workflow, not for evidence of adherence, which the video's description does not claim.

What each one encodes well, and what it loses

The useful question is not “which is smarter” but “what information does each format carry cheaply.” The table separates what the specs allow from what teams usually do.

What each artifact carries in practice (inference, based on the two specs)
InformationDTCG JSON tokensDESIGN.md
Exact values and typesCore purpose; typed and validated by toolingYes, in YAML front matter; tokens are normative
Aliases (semantic names pointing at base values)Yes, {group.token} and JSON PointerYes, {path.to.token} references; a lint rule flags broken ones
Multi-platform outputStrong: designed as input to build toolsIndirect: export to Tailwind or DTCG via the CLI
Where a value applies (scope)Possible only through descriptions or custom extensionsNatural in prose sections
Why a pairing exists (contrast intent)Rarely captured; names are the only hintNatural in prose; a contrast-ratio lint exists
Prohibitions (never use X for Y)No standard placeA dedicated Do's and Don'ts section
Diff and review in version controlClean, line-based, machine-checkableProse diffs are harder to check automatically

Two points matter for the rest of the argument. First, JSON can hold prose: the DTCG format supports descriptions on tokens, and extensions for tool-specific data, so rationale does not strictly need Markdown. Most real token files simply leave them empty. Second, a DESIGN.md can fail in the opposite direction: prose that contradicts the tokens. The spec settles that conflict in favour of the tokens, but an agent has to follow the rule.

Two-column comparison of what DTCG JSON tokens encode versus what Markdown prose encodes, with a bracket noting names and values overlap.
The overlap is names and values; the difference is rules, scope and rationale, which JSON can hold but usually does not.

The evidence ladder: what exists, and how strong it is

Searching for a controlled comparison of design-token formats for agent adherence turned up no peer-reviewed study and no public reproducible benchmark that isolates the format. That absence is a finding, and it comes with a caveat: absence from the sources checked on October 1, 2026 is not proof that none exists privately. What does exist falls into four rungs.

Evidence relevant to the question, ranked by how directly it tests it
SourceWhat was actually doneLabelLimit
Open Design post (May 23, 2026)Three agent runs per condition, JSON tokens only versus DESIGN.md, one landing-page hero, one modelReported (vendor)Judged by eye; no raw data, metric or steps to reproduce
WaveSpeed post (May 15, 2026)One author rebuilt a component library with DTCG-style tokens and with DESIGN.md; eight screens over two daysReported (individual)Qualitative; no counts; the author notes it is not a paradigm shift
Indeed MCP benchmarkEight MCP configurations over one knowledge base, comparing metadata formats for component documentation retrievalReported (practitioner)Measures retrieval, not token adherence; prompt count and scoring method not given on the page
CHI 2026 extended abstractCompared instruction-based, context-based and registry-based strategies on six real UIsDocumented (peer-reviewed abstract)Not a Markdown-versus-JSON comparison; abstract only was read
He et al., 2024 and Kato and Kato, 2026Same content in plain text, Markdown, JSON, YAML and other formats on non-design tasksDocumented (preprints)Different tasks and models; no design tokens
Evidence ladder with vendor anecdotes at the bottom, practitioner write-ups, studies on other tasks, and an empty top rung for controlled design-token studies.
The top rung, a controlled study of design-token formats, is empty in the sources checked.

Reading the Open Design claim closely

The Open Design article, written by the opendesigner.io editors, is the most direct claim that Markdown wins. It describes its test this way: same brand brief, same target artifact (a landing-page hero), three agent runs each against a JSON-tokens-only input and against a DESIGN.md input, using Claude Opus 4.7. It reports that with JSON only the agent used tokens correctly “but didn’t seem to understand why,” and that DESIGN.md output was visibly closer to the brand intent.

That is a legitimate observation and a fair hypothesis. It is not a measurement, for four reasons visible on the page.

  1. Format and information are confounded. The DESIGN.md condition contains prose rules the tokens-only condition lacks. A better result may come from the extra rules, not from Markdown. The test cannot separate the two.
  2. The outcome is a judgment. “Visibly closer” has no metric, rubric or blind rater described.
  3. The sample is tiny and not reconciled. Three runs per condition is six runs; the same page also mentions “8 out of 10 agent runs we audited.” The page does not explain how those numbers relate.
  4. The publisher has an interest. The post is part of a project that promotes DESIGN.md-style design systems. That does not make it wrong, but it is a reason to wait for independent replication.

The same pattern repeats in other write-ups. A site that sells a library of DESIGN.md files argues the formats are complementary and describes a pipeline, but offers no test. A practitioner post argues machine-readable JSON or YAML reduces hallucination and cites a single dashboard project. None of these is evidence against its author's good faith; none is a comparison you could rerun.

What the adjacent research does and does not tell you

Studies of prompt format exist, but on other tasks. He et al. (arXiv 2411.10541, submitted November 2024) gave the same content to OpenAI GPT models as plain text, Markdown, JSON and YAML. Performance of GPT-3.5-turbo varied by up to 40% on a code translation task depending on the template, while GPT-4 was more robust, and the authors conclude there is no universally optimal format. Kato and Kato (2026) varied the format of algorithm specifications across seven styles and three models over 4,020 generated implementations; the best format differed by model, and one model showed no format differences once the specification was complete.

Taken together, these support a cautious inference: format effects are real for some models and tasks, shrink as models get stronger or the content gets more complete, and do not point to one winner. Neither study used design tokens, a UI task, or a modern coding agent that can read files, run linters and iterate, so they cannot be projected onto your case.

The two nearest design-system results answer different questions. The Indeed benchmark compared how component documentation should be stored for retrieval through an MCP server (Model Context Protocol, the interface agents use to call external tools and data). Its page defines JSON as structured component metadata and Markdown as human documentation, and its table reports its best configuration at 88% token savings and 92% retrieval accuracy against an in-production baseline at 89%. That is about getting the right component facts into context cheaply. It says nothing about whether an agent then follows a color rule. Note also that a secondary article repeats different figures for this benchmark that do not appear on the primary page; this article uses only the primary page.

The CHI 2026 extended abstract is the closest thing to a controlled design-system study. It compared three ways of giving an agent a design system and reports the highest compliance (95.08%) for the registry-based strategy, where the agent assembles pre-built components, across six real-world UIs. Its message, in the abstract, is that pre-built components beat style guides in the prompt. If that holds up, the larger lever for adherence is giving the agent components to compose, not choosing between two ways of writing down colors. Only the abstract was read for this article, so treat it as a pointer to the full paper.

A protocol to test adherence on your own codebase

Because nobody else has done the clean version, the cheapest route to an answer is running your own. The design below is a proposal (inference), sized for a team that can spare an afternoon of compute and an hour of review.

  1. Fix the information first. Write your design rules once, in plain statements: which token for which role, scope rules, contrast constraints, prohibitions. Every condition below must contain exactly these rules, so that format is the only variable.
  2. Build four conditions. A: values only, as DTCG JSON. B: the same JSON with the rules placed in $description fields. C: a DESIGN.md with the same tokens and the same rules in prose. D: the hybrid in the next section. Keep token names identical across all four.
  3. Pick five to eight tasks from real work, from easy (restyle a button) to hard (a new settings page with error and empty states). Fix the model, version and settings, and record them.
  4. Run each task in each condition at least five times, in randomized order, from a clean checkout so earlier output cannot leak.
  5. Score with machines first. Search the generated code for literal hex colors, pixel sizes and font names that are not in the token set (hard-coded values). Run your linter and, for DESIGN.md, the project's own lint for broken references.
  6. Score rules with a checklist. For each prose rule, mark followed or violated per output. Have someone who does not know the condition do it, with file names stripped.
  7. Report counts, not just averages: violations per output per condition, plus tokens consumed. With five runs you can only trust large differences; say so.
Flow of the adherence test: fix the rules, build four conditions, run each task five times, lint for hard-coded values, checklist for rules, compare violation counts.
Hold the information constant, vary only the format, and count violations per output.

Condition B is the one the published comparisons skip. If A loses to C but B matches C, the format was never the problem; the missing rules were. If C beats B, then format or structure matters for your model, and you have your own evidence for it.

A pragmatic hybrid (opinion)

This section is opinion, built on the documented behaviour above rather than on adherence data. If you already run a token pipeline, keep the DTCG JSON as the single source of truth for values. It is what Style Dictionary, Terrazzo and Figma-side tooling consume, and the 2025.10 report is stable, with the tooling caveats noted earlier.

Generate the token block of a DESIGN.md from that source rather than maintaining values twice, then hand-write only the prose: scope rules, contrast intent, prohibitions and edge cases. The DESIGN.md tool documents export to the DTCG format, not the reverse, so the JSON-to-front-matter step is a small script you would write yourself. Add a CI check that every value in the front matter equals the JSON value, so the two cannot drift. Because DESIGN.md is alpha, isolate it behind that script and expect to adjust.

For agents that call an MCP server, keep component APIs (props, variants, sizes) as structured JSON that the server returns on request, and keep behavioural rules in Markdown. That split matches the Indeed team's stated conclusion, though the benchmark behind it measured retrieval rather than adherence.

Finally, consider whether a registry of real components, as in the CHI abstract, would remove more mistakes than any file format. An agent that composes Button from your library cannot invent a hex value for it.

Pick by job, then measure the format

Markdown or JSON is a weaker question than it looks. JSON in the DTCG format is the better home for values that tools must validate and compile. Markdown is the better home for the reasons, scopes and prohibitions that raw tokens rarely carry. DESIGN.md puts both in one file and says plainly that tokens win when they conflict with prose.

Whether agents follow the prose-plus-tokens version more reliably than tokens alone has not been shown in a controlled, public, reproducible way. The strongest published claim comes from a small test with a confound, a judged outcome and a publisher with a stake. Treat it as a hypothesis worth testing, and run the four-condition protocol before you rewrite your design system documentation around it. Specs, tooling and models are changing quickly, so recheck the sources linked below before you rely on any of this.