zig-Hocon is a HOCON (Human-Optimized Config Object Notation) parser for Zig.
Why? I have been using HOCON in production for nearly five years, in other languages. It is the format I reach for when one service runs in several environments: a shared file is included once, each environment overrides only the keys that differ, substitutions pull one value into another, and the merge rules decide the rest — no templating layer on top.
What this library adds is the other half: the schema stays in the language you
already work in. You describe the config as an ordinary Zig struct, HOCON is
only what sits on disk, and the parser fills the struct in — the same shape as
std.json.
Requires Zig 0.16.0 or newer.
Status: early development. Source text parses into a value graph, and the graph converts into a Zig struct of your own through
hocon.parseFromSlice, shaped afterstd.json, or prints as JSON from thehoconcommand line tool. Substitutions are not resolved yet — see Features below for what currently works, and Conformance for how far that gets. No tagged release yet: install straight from the repository, as below.
Add the package to your build.zig.zon:
zig fetch --save git+https://github.com/lejmr/zig-hoconThat records the dependency in build.zig.zon, pinned by its content hash. For
a particular commit, append it: git+https://github.com/lejmr/zig-hocon#<commit>.
Then hand the module to your executable in build.zig:
const hocon = b.dependency("hocon", .{ .target = target, .optimize = optimize });
exe.root_module.addImport("hocon", hocon.module("hocon"));and @import("hocon") works in your code.
Describe the config as a struct and parse into it, the way std.json does:
const hocon = @import("hocon");
const Server = struct {
host: []const u8,
port: u16 = 8080,
tls: ?struct { cert: []const u8 } = null,
level: enum { debug, info } = .info,
};
const parsed = try hocon.parseFromSlice(Server, gpa,
\\host = example.org
\\level = debug
);
defer parsed.deinit();
const server = parsed.value; // server.port == 8080, server.tls == nullparsed owns an arena with everything the result points to, and deinit gives
it back in one go — any allocator will do. Strings in the result are copies, so
the source can be freed as soon as the call returns. When you already have an
arena of your own, parseFromSliceLeaky(Server, arena, source) returns the
plain struct and allocates straight into it.
Reading the file is yours for now. Zig 0.16 hands back the 0-terminated buffer
the parser takes when asked for a 0 sentinel, so the whole thing is:
const source = try std.Io.Dir.cwd().readFileAllocOptions(
io, "server.conf", gpa, .limited(1 << 20), .of(u8), 0,
);
defer gpa.free(source);
const parsed = try hocon.parseFromSlice(Server, gpa, source);
defer parsed.deinit();
const server = parsed.value;For the tree itself rather than a struct, ask for hocon.Value, the way
std.json.Value works; std.json.Stringify prints it:
const parsed = try hocon.parseFromSlice(hocon.Value, gpa, source);
defer parsed.deinit();
try std.json.Stringify.value(parsed.value, .{}, writer);zig build also installs hocon, which prints a file as compact JSON:
$ zig build
$ zig-out/bin/hocon server.conf
{"server":{"host":"example.org","port":8080,"tags":["a","b c"],"tls":"on"}}Unquoted true, false, null and numbers keep their JSON type, a number is
printed as written (1e5 stays 1e5), and anything quoted is a string. A file
that does not parse, or still holds a substitution, exits 1 with the reason on
stderr. zig build run -- server.conf does the same without the install step.
Tracks coverage of the HOCON spec, following its own section names so the two can be read side by side. Checked items are implemented and tested; unchecked items are planned.
Everything under §Syntax, grouped in the order it is meant to be built rather than the spec's, and blank lines separate the stages. See Order of work below for why they fall this way.
- Unchanged from JSON — a JSON document parses as HOCON
- Comments (
#and//) - Omit root braces
— both
a = band{ a = b }are a valid root - Key-value separator
—
=and:are interchangeable, and may be omitted before{ - Commas — interchangeable with newlines, runs collapse, a trailing one is allowed
- Whitespace — kept verbatim inside a value, trimmed at its ends
- Unquoted strings — including the reserved characters that end one
Not spec sections of their own, but needed to get there, and already done:
-
Quoted and unquoted values stay distinguishable in the tree, which type inference needs to tell
a = "1"froma = 1 -
A value with nothing in it (
a =) is a parse error rather than a crash -
Multi-line strings — a text part like any other, on the key side and as an include path too; the
"""stay in the tree because escapes are literal inside them, which is what tells"""a\nb"""from"a\nb" -
Includes (file, url, classpath, required)
-
Path expressions — a substitution path is split into elements when its reference is built, so resolution never parses. The split walks the parts the parser left rather than the text, which is what keeps a quoted dot out of it: a dot outside quotes is the only thing that ends an element, and the end of a part means nothing at all. So
${x."y.z"}is two elements and finds the keyx { "y.z" = 1 },${a"b"}is the single elementab, and${a."".b}keeps its empty middle, which java accepts. Still to do: rejecting.a,a.anda..basBadPath, and lending the same splitter to the key side -
Paths as keys — an unquoted dotted key nests:
a.b.c = 1builds the same tree asa { b { c = 1 } },a.b {c = 1}joins the two ways of writing it, and a number splits like any other unquoted string, so1.5 = xis1 { 5 = x }. A malformed path (.a,a.,a..b) is rejected, though as the parser'sUnexpectedTokenrather than theBadPathjava names. Known gap: a key written in several parts is left alone, so"a"."b" = 1anda."b.c" = 1stay flat where java nests them. The splitter that path expressions use answers exactly this and only wants wiring up here -
Duplicate keys and object merging — a key written twice is one member, whether the two spellings meet inside one block (
{b=1, b=2}) or in a concatenation ({b=1} {b=2}), and the surviving member keeps the position of the first. A key held by both sides merges again only if both hold an object; anything else lets the right-hand value win outright, so{b=[1]} {b=[2]}is[2]where the[1] [2]of a concatenation would be[1,2]. That pair is the whole difference between merging and concatenating -
Value concatenation — the parser collects the parts of
a = x "y",a = [1] [2]anda = {x=1} {y=2}with the whitespace between them, and the value graph then joins them: lists concatenate flatly, objects merge, text parts join and stop being numbers. Mixing kinds is theWrongTypejava reports, and it is reported while loading rather than later. The whitespace between two parts goes whichever way its neighbours do — a separator beside a list or an object, a character beside text — and a quoted space is a value, so[1] " " [2]is an error where[1] [2]is[1,2]. A substitution is the one part that cannot be joined into, so the parts around it are kept in order for resolution to finish -
Substitutions (
${a.b.c},${?a.b.c}) — parsed, not resolved:${…}and${?…}become nodes of their own wherever a value may stand, and everything java rejects at parse time is rejected here too (${}, an unclosed${a, a newline or a nested${inside the path,?anywhere but directly after${, and a substitution where a key or an include target belongs). In the value graph a substitution becomes a reference of its own and everything holding one stays pending, since its type is unknown until it resolves — which is why"1" ${x} [2]loads and only fails oncexturns out to be a number, while"1" [2] ${x}fails immediately. The path is split on.there, but naively: a quoted dot (${a."b.c"}) still splits and waits for path expressions above -
The
+=field separator — self-referential array append, which the spec defines as sugar fora = ${?a} [b];+is rejected by the tokenizer today
Order of work — why the stages above fall this way
The guiding idea is to build one complete tree first and only then walk it. That is what puts includes so early: until they are spliced in, there is no complete tree to walk, because an include contributes members the merge would otherwise never see.
Two constraints from the spec fix the rest of the order. An include has to be resolved before anything reads the values it contributes, and substitutions resolve against the finished object graph. Getting the second one wrong is a known pyhocon bug and one of the reasons this project exists.
Includes next, as a tree operation: parse the referenced file and splice its root members in. Nothing here needs evaluation, which is why it can come this early, and it has to precede substitutions regardless.
Path expressions and paths as keys before merging, not after. Merging has to
know that a.b = 1 and a { b = 2 } describe the same field, and a substitution
written as ${a.b} looks it up the same way — so splitting a key into segments is
part of building the graph rather than something bolted on later.
Duplicate keys, merging and value concatenation turned out to belong to
building the value graph rather than to the walk over it. Java decides them while
parsing too — a = {x=1} {y=2} prints merged without anything being resolved —
and doing the same here buys an invariant worth having: once loading is done,
every value that holds no substitution is finished. Resolution then never has to
know what a quote or a gap meant, only how to put a value where a reference was
and re-run the same two functions loading used.
Substitutions are what is left for the walk, and the reason a value holding
one stays pending: its type is unknown until the graph is finished, so the type
check that rejects 1 [2] cannot run across it. += then costs almost nothing,
since the spec defines a += b as sugar for a = ${?a} [b].
Somewhere around here the library becomes useful end to end: with merging done and
automatic type conversions in place, a JSON rendering works — and type
conversion is on the critical path rather than a convenience, since rendering has
to decide whether a = 1 is the number 1 or the string "1".
Conversion of numerically-indexed objects to arrays last of the syntax items: a small rule with few users.
§API Recommendations is explicitly advisory — a HOCON implementation is conforming without it — but these are what makes the format pleasant, so they are on the roadmap.
- Automatic type conversions
— into a Zig type rather than through typed getters: integers, floats,
bool(true/yes/on,false/no/off, lowercase only), strings, slices, structs, optionals, enums and field defaults. A number reads as a string exactly as written (1.50stays"1.50"), unquotednullfits only an optional, and objects and arrays never become strings. One deliberate difference from typesafe/config:1.5into an integer is an error rather than a silent1. Not yet: numerically-indexed objects as lists. - Duration format (
10s,5m, ...) - Period format (
3d,2 weeks, ...) - Size in bytes format (
512K,1G, ...) - Config object merging and file merging (
with_fallback-style) - Substitution fallback to environment variables
Deliberately out of scope: Java properties mapping, conventional JVM config file names, and override by system properties — all of them describe JVM conventions rather than the format.
- Tokenizer — objects, arrays, unquoted, quoted and triple-quoted strings,
comments, the reserved characters that terminate an unquoted string, and
the substitution openers
${and${?. Numbers, booleans andnullare tokenized as plain strings; giving them types belongs to evaluation. What is inside${…}gets no special lexing — it is an ordinary token stream closed by}. A token'sloccovers the token, quotes included, rather than its content: the tree keeps the raw text, so anything else would make every reader add the quotes back — one of them, or three. Known gaps: an escaped backslash right before a closing quote (\\"), and a leading+on a number. - Parser — source text to a syntax tree, covering the checked syntax items above. Nested objects and arrays to any depth. The tree keeps whatever distinguishes two sources that mean different things — quotes, the whitespace between two parts of a value, duplicate keys — and leaves what it means to evaluation.
- Memory — one arena owns the whole document, and everything the parser and
the value graph produce is allocated from it or borrowed from the source
text. Nothing inside has a
deinitof its own: freeing is dropping the arena.Parsed(T)is that arena packaged with the result, so its singledeinitis the only one a caller sees; the arena lives behind a pointer, which keeps any copy ofParsedfreeing the same memory. The input text must outlive the value graph, since its leaf values are slices into it — a converted struct does not have that tie, its strings are copied. TheLeakyvariants skip the packaging and expect an arena from the caller: an individual string is not freeable on its own, so any other allocator leaks by construction. A failed parse, out of memory at any allocation included, gives everything back. - Value graph — syntax tree to values. A key is a type rather than a
string, because three places produce one and all three have to agree that
a,"a"and"""a"""are the same key; unquoting is shared with values, where the delimiter is kept as a flag instead, sincea = "1"anda = 1differ only by it. Concatenation, object merging and path splitting happen here (see the syntax items above), which leaves a value either finished or explicitly pending on a substitution. The rule the layer is built on: anything decidable while the context is still around gets decided now, because nothing downstream can reconstruct it — which is why a gap beside a list is dropped and a gap between two references is not. - Conversion — value graph to a caller's type, one
switchon@typeInfo(T)with a branch per kind of type, recursing through struct fields, slice elements and optionals. Values stay text until a type asks for them, soyesis a boolean only where aboolis wanted. A key missing from the file takes the field's default, thennullfor an optional, and only then is an error. A substitution still pending is an error rather than a guess, until evaluation exists. - Evaluation — resolving substitutions against the finished graph and splicing includes.
- Public API, first cut —
hocon.parseFromSlicereturningParsed(T)and itsparseFromSliceLeakytwin, named and shaped afterstd.json, withhocon.Valueas a target for the tree itself. - JSON rendering —
Value.jsonStringify, whichstd.json.Stringifyfinds on its own, so there is no serializer here to maintain. A scalar is text until this point; quotes decide string, and otherwise JSON's own number grammar decides number, all of the text or nothing (.5,1.and1 2stay strings). A tree with a substitution left in it is refused when it is parsed, since Stringify has no way to report an error of its own. This is the bridge the project is ultimately for. - Command line —
hocon <file.conf>prints JSON, exits 1 on rejection; the same contract as a conformance adapter, so it is what the suite measures (conformance/adapters/zig-hocon.sh). - Public API, the rest —
ParseOptions(ignore_unknown_fields; unknown keys are ignored for now),parseFromValue, andparseFromFiletaking anstd.Ioand a directory, whichincludewill need anyway.
Disputed behaviour is settled by running it, not by reading the spec from
memory. tools/oracle/ holds two reference implementations behind one interface:
printf 'a = grumpy wombat\n' | tools/oracle/hocon-java # Lightbend typesafe/config
printf 'a = grumpy wombat\n' | tools/oracle/hocon-py # pyhoconhocon-java is the authority; hocon-py shows what the implementation this
project replaces would do. See tools/oracle/README.md
for the input format and the list of confirmed divergences between the two.
To know what "compatible" means, conformance/ turns the specification into
data: every normative sentence of HOCON.md becomes a config file and the value
it must produce. Expected values are seeded from typesafe/config, then checked
by hand against the sentence they stand for; where the two disagree the spec
wins and the row says so. Where it stands today:
| implementation | against the spec | against typesafe/config |
|---|---|---|
| Java typesafe/config 1.4.9 | 89% (390/436) | 99% (433/436) |
| pyhocon pyhocon 0.3.63 | 71% (308/436) | 67% (291/436) |
| zig-hocon zig-hocon code@136b586 | 57% (248/436) | 59% (258/436) |
As you can see there is still a lot of work ahead, but I believe it is worth every second.
A set of conformance tests on top of that keeps the implementation clean and fully working — every case, and what each implementation printed, is in the conformance report.
The suite is not tied to Zig. Any HOCON parser can be scored on it through a ten-line adapter — conformance/README.md shows how to add yours.
This project is developed in the open. If you need a HOCON feature that isn't implemented yet, please open a PR — feature requests as PRs (even a failing test showing what you need) are the fastest way to get something prioritized. See CONTRIBUTING.md for the PR workflow (draft until CI is green).
BSD 3-Clause — see LICENSE.
