A parser for the DOT graph language, the text format used by Graphviz, written in Zig.
It reads DOT text and gives you its structure: nodes, edges, attributes and subgraphs. It also gives clear error messages. It does not draw graphs or compute layouts. You decide what to build on top, for example a linter, a formatter, or a converter into your own graph type.
The package has two independent modules:
dot_parserparses DOT.markup_parserparses HTML-like markup, such as Graphviz'slabel=<<b>Hi</b>>labels. It can check labels while DOT is parsed, or work on its own with no DOT at all.
Status: experimental (0.x). It works and is well tested, but names may still change between versions. Breaking changes are listed in the changelog.
digraph Pipeline {
node [shape=box];
fetch -> parse -> check [color=red];
subgraph cluster_output { render; save }
check -> { render save }
}- Helpful errors. Each problem points at the exact spot, explains what is wrong, and often suggests a fix. Parsing keeps going after an error, so you see more than one problem at a time.
- You choose where memory comes from. Use any allocator, an arena, or fixed buffers with no heap allocation at all.
- Runs anywhere. No dependencies, no OS calls, no file access. It builds for embedded targets and WebAssembly.
- Keeps what was written. Spelling, order and duplicates are kept exactly. Nothing is silently rewritten or interpreted.
- Supports safe handling of untrusted input. Parsing uses no recursion and makes one pass over the input. Limits are off by default, so you set explicit budgets for size, nesting and memory, and can parse step by step with cancellation.
- Configurable. Strict by default, with a lenient mode and optional extra checks.
- HTML-like markup, inside DOT or on its own. Check the inside of labels
such as
label=<<b>Hi</b>>while parsing DOT, or parse HTML-like markup by itself (development version only for now). - Bring your own processor. Swap in your own label checker at compile time. Its settings plug into the same settings system as DOT's.
const std = @import("std");
const dot = @import("dot_parser");
pub fn main(init: std.process.Init) !void {
const allocator = init.gpa;
var buffer: [4096]u8 = undefined;
var stdout_writer: std.Io.File.Writer = .init(.stdout(), init.io, &buffer);
const stdout = &stdout_writer.interface;
defer stdout.flush() catch {};
const source =
\\digraph {
\\ a -> b;
\\ b -> c [color=red];
\\}
;
// Problems are collected here instead of being printed or thrown.
var bag = dot.GrowableDiagnosticBag.init(allocator, .{});
defer bag.deinit();
// Parse the text, then check it (for example, `--` inside a digraph).
var result = dot.parseAndValidate(allocator, source, bag.sink(), .{});
defer result.deinit(allocator);
if (!result.documentValid()) {
for (bag.items(), 1..) |problem, number| {
try dot.console.renderBoxed(problem, number, .{ .source = source, .source_name = "graph.dot" }, stdout);
}
return;
}
const document = result.document.?;
var edges = document.edgeIterator();
while (edges.next()) |edge| {
// An edge end is a node or a whole subgraph (`a -> { b c }`).
if (edge.left != .node or edge.right != .node) continue;
const from = document.nodeReference(edge.left.node).?.identifier;
const to = document.nodeReference(edge.right.node).?.identifier;
try stdout.print("{s} {s} {s}\n", .{
document.text(from), edge.operator.lexeme(), document.text(to),
});
}
}This prints:
a -> b
b -> c
If line 3 said b -- c instead (an undirected edge in a directed graph), you
would get:
┌─ Error 1: edge operator does not match the graph kind
│ graph.dot:3:7
│
│ 1 │ digraph {
│ │ ─────── the document is directed because of this keyword
│ ⋯
│ 3 │ b -- c [color=red];
│ │ ^^ expected '->', found '--'
│
│ Hint: change '--' to '->', or declare the document with 'graph'
│ Fix: replace '--' with '->' (one possible repair)
└─ E1 ─ [dot_parser:E.Validation.Operator.002]
This program is examples/quick_start.zig. Getting started walks through it step by step.
You need Zig 0.16.0. Add the package to your project:
zig fetch --save git+https://github.com/AshutoshMahala/dot-parser#mainThen import the modules you need in your build.zig:
const dot_parser = b.dependency("dot_parser", .{ .target = target, .optimize = optimize });
exe.root_module.addImport("dot_parser", dot_parser.module("dot_parser"));
// Only if you check labels or parse markup:
exe.root_module.addImport("markup_parser", dot_parser.module("markup_parser"));This installs the development version, which is what these docs and examples
describe. The latest release, 0.3.0, has an older API: for example, it has no
GrowableDiagnosticBag and no markup_parser. If you need it, use #v0.3.0
and follow the README at that tag.
The differences are listed under "Unreleased" in the changelog.
All of DOT's statement syntax:
graphanddigraphdocuments, with optionalstrictand a name- Nodes, edges, and edge chains (
a -> b -> c) - Attributes:
[color=red]lists,node [shape=box]defaults, andrankdir=LR - Subgraphs, named or not, nested, and as edge ends (
a -> { b c }) - Ports (
a:out,a:out:n) - Every kind of name: plain words (including non-ASCII), numbers,
"quoted strings"joined with+, and HTML-like<...>labels - Comments (
//,/* */,#) and optional semicolons
See supported syntax for the full list and the few places where it differs from Graphviz.
It reports what the file says. It does not work out what the file means:
- It doesn't lay out or draw anything.
- It doesn't interpret attribute values.
color=redis just two pieces of text. - It doesn't apply defaults.
node [shape=box]stays a statement; it isn't copied onto each node. - It doesn't expand
a -> { b c }into two edges, merge repeated subgraphs, or build a list of unique nodes. - It doesn't read files. You pass it bytes.
- It doesn't keep comments, so it can't reformat a file on its own.
You can do all of these on top of the parsed document. Why it works this way explains these choices.
DOT
| Guide | Read it when you want to… |
|---|---|
| Getting started | parse your first DOT file and read the result |
| Reading a parsed graph | walk nodes, edges, attributes, ports and subgraphs |
| Supported syntax | check exactly which DOT input is accepted |
| DOT error codes | look up a DOT error or warning |
Markup
| Guide | Read it when you want to… |
|---|---|
| Checking HTML-like labels | check <...> labels inside DOT files |
| Parsing markup on its own | parse HTML-like markup without DOT, or check many fragments |
| Bringing your own processor | plug your own label checker into DOT parsing |
| Markup error codes | look up a markup error or warning |
Both parsers
| Guide | Read it when you want to… |
|---|---|
| Errors and diagnostics | show errors, understand results, or apply suggested fixes |
| Memory | use an arena or fixed buffers, or work with no allocator at all |
| Settings | make parsing stricter or more lenient, add checks, or set limits |
| Parsing in small steps | spread parsing over time, or cancel it |
Project
| Guide | Read it when you want to… |
|---|---|
| Why it works this way | understand the main design decisions |
| Roadmap | see what is planned but not built yet |
| Performance | see measured speed and memory use |
| Architecture | find your way around the source code |
The same index is in docs/README.md.
Small runnable programs in examples/. zig build examples builds
and runs them all, and installs each one in zig-out/bin/.
| Example | Shows how to… | Guide |
|---|---|---|
| quick_start | parse, show problems, and list edges | Getting started |
| parse_undigraph | print every kind of statement | Getting started |
| identifiers | get the real value of a quoted or joined name | Reading |
| attributes | read attribute lists, defaults and assignments | Reading |
| edge_chains | walk chains like a -> b -> c |
Reading |
| ports | read ports like a:out:n |
Reading |
| subgraphs | walk subgraphs and their contents | Reading |
| subgraph_endpoints | handle edges that end at a subgraph | Reading |
| diagnostics_demo | print problems with source lines and color | Errors |
| check_file | build a command-line checker for .dot files |
Errors |
| fixed_buffer | parse with no allocator at all | Memory |
| policies | use checks, lenient mode, graph kinds and run-time settings | Settings |
| bounded | parse in small steps, and cancel | Small steps |
| composed_markup | check every label while parsing DOT | Labels |
| delayed_markup | check only the labels you choose | Labels |
| markup | parse and validate markup on its own | Markup |
| custom_processor | plug in your own label checker | Own processor |
zig build test # all tests
zig build examples # build and run every example
zig build check-freestanding # compile for RISC-V32 and Wasm32
zig build bench -Doptimize=ReleaseFast # main benchmarkLicensed under either of
- MIT license (LICENSE-MIT)
- Apache License, Version 2.0 (LICENSE-APACHE)
at your option (MIT OR Apache-2.0).
The optional console renderer uses Unicode-derived width tables under the Unicode License v3. The full notice is in console_widths.zig.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this work by you shall be dual licensed as above, without any additional terms or conditions.