← back to home

Zinc String Interpolation

· 869 words · 5 min read

I wanted a simple version of string interpolation. The one every language has:

"hello {name}, your age is {age}"

No new operators, no clever syntax or mini-language embedded inside a string. That's the design of a lot of languages like javascript, c++ or python and others and this is the entire spec.

But in Zinc is slightly different because it must care about memory allocation and how this interpolated string is compiled.

The problem is that there's no allocator to reach for: to turn {age} into text, something has to produce bytes. age is an i32; its decimal form is somewhere between one and eleven characters long. Those characters then have to be joined with the literal runs around them, In most languages this is invisible because there's a global allocator.

Zinc doesn't have one. Allocation is explicit. You have to pass an allocator or you can't allocate. That constraint is the point of the language and string interpolation is exsactly the kind of convenience feature that quietly breaks it.

Then what I thought at first is: Zinc does have a way to get an allocator without threading it through every signature: The capability system. You write a with a: Allocator and the compiler finds it in the scope. So interpolation could desugar into code that chain the string with it.

For this reason the compiler has to know what Allocator is. Currently is just a user-defined type but with this approach it needs to became a lang-item: a type the compiler recognises by role rather than by name. I didn't want to make it a lang-item because it binds the compiler to a memory model. The moment the compiler emits a call to alloc, it has opinions, about who frees, about what happens when allocation fails and about how long the result lives. Those opinions leak into every program, whether or not it ever interpolates a string. Zinc's whole position is that the language shouldn't holde those opinions. I just want to keep the Allocator a type defined in the standard library and treat it as a normal type.

Another problem is that the result is not always a string: Even setting the allocator aside, building a string is the wrong answer most of the time. If you want to print a string to the stdout or log into a file or even send bytes to a socket you don't need to allocate anything. The only case you have to allocate is when you want to build a string and use it. SO interpolation shouldn't return anything. It should go somewhere.

The first solution I had was defining a facet for "things that accepts bytes", called Writer and let the compiler treat a writer applied to an interpolated literal as a special form.

stdout := Writer(1)
stdout("hello {name}, your age is {age}")

This fixes the destination problem, and I kept the Writer facet from it. But I don't like the special case where a value is callable only for a specific type. And in this case the interpolated literal doesn't have its own value: it becomes a value only when it's fired by the writer. The other problem is that a callable writer must have always the same return type and it couldn't return always void: to build a string i need to return the allocated string.

The final solution is even simpler: the compiler compiles and interpolated string as an array of objects that represent the interpolation:

#lang(interpolated_string)
pub InterpolatedString :: enum {
	Literal([]char)
	Obj(Writable)
}

A segment is either a run of literal text or a value that knows how to write itself.

"hello {name}, your age is {age}"
will compiled as [4]InterpolatedString:
[ Literal("hello "), Obj(name), Literal("your age is "), Obj(age) ]

It has a type, it's a value and you can pass it to a function and the function decides what to do with it. This design introduced 3 different lang-items:

#lang(writer)
pub Writer :: facet {
	write: fn (*u8, usize) isize
}

#lang(writable)
Writable :: facet {
	sink: fn (Writer) u0
}

and the InterpolatedString declared above.

Then a function to handle this list becomes really straightforward:

pub interpolate :: fn (w: Writer, items: []InterpolatedString) {
	for item in items do match item {
		.Literal(chars) -> w.write(chars.ptr, chars.len),
		.Obj(obj)       -> obj.sink(w)
	}
}

No format-string parsing, no intermediate buffer, no varargs. And the print function is a few lines of code:

print :: fn (items: []InterpolatedString) {
	stdout := Writer(FWWriter{ fd = 1 }.&)
	interpolate(stdout, items)
}

So the compiler work was not too much: it splits the literal into alternating literal runs and holes. Resolve each hole's expression and check it satisfies Writable and resolves the literal to []InterpolatedString. Such that the compiler never emits weird calls and never knows what a file descriptor is or the memory model of the language.

There are some things that are not yet supported:


← back to home · rss