<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Radu&apos;s little blog</title><description>My little place on the internet :^)</description><link>https://raung0.github.io/</link><item><title>Optimizing pipelines through specialization</title><link>https://raung0.github.io/blog/optimizing-pipelines-through-specialization/</link><guid isPermaLink="true">https://raung0.github.io/blog/optimizing-pipelines-through-specialization/</guid><description>Optimizing pipelines through specialization</description><pubDate>Thu, 03 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;When designing how my video editor &lt;a href=&quot;https://github.com/raung0/kama_studio&quot;&gt;Kama
Studio&lt;/a&gt; should work, I had one very
important requirement in mind: &lt;strong&gt;previews should be instant&lt;/strong&gt;.  Considering that
I wanted to be able to edit even 4K 60FPS video, it was clear to me that I
needed to utilize the GPU &lt;em&gt;somehow&lt;/em&gt; to handle complex effect pipelines.&lt;/p&gt;
&lt;p&gt;While thinking about it whilst I was doing some random bs throughout the house,
I came up with an idea: “What if I could stitch some shaders together and have
all effects run that way?”.  Welp, so I did.  I decided to make pretty much all
video effect shaders, through the power of &lt;a href=&quot;https://lib.rs/crates/naga&quot;&gt;naga&lt;/a&gt;,
able to achieve that goal.  Each plugin and effect is namespaced, so there are
no conflicts, and they all get stitched together into a single combined shader
at the end.  This is what allows my video editor to reach the performance it
does, and as such allows the shader compiler to not only better optimize the
code, but also reduce memory bandwidth and needed passes.&lt;/p&gt;
&lt;p&gt;So I was thinking, what if you could apply this to other programs?  Turns out,
you can!&lt;/p&gt;
&lt;p&gt;By using various backend compiler technologies, such as &lt;a href=&quot;https://llvm.org&quot;&gt;LLVM&lt;/a&gt;
or even
&lt;a href=&quot;https://github.com/bytecodealliance/wasmtime/tree/main/cranelift&quot;&gt;cranelift&lt;/a&gt;,
you can generate some code on the fly and make things run significantly faster.
Browsers do something like this through &lt;span title=&quot;Just-in-Time&quot;&gt;JIT&lt;/span&gt;,
where JavaScript is compiled when needed, allowing you to play browser games
with decent performance.  If you have a deterministic data structure, like Kama
Studio’s effect graph for instance, you already know how everything should be
hooked up together and the possible values you can have, so you can emit code
representing that entire pipeline.  Many programs can benefit from this, for
example &lt;span title=&quot;Digital Audio Workstation&quot;&gt;DAWs&lt;/span&gt;, where you can chain
different effects on top of one another and then even make the backend compile
to the user’s CPU’s specific features while also avoiding expensive operations
that come from a generalized architecture.&lt;/p&gt;
&lt;p&gt;Let’s say that we have a basic effect chain for a microphone:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Input -&amp;gt; Noise Reduction -&amp;gt; EQ -&amp;gt; Limiter -&amp;gt; Output
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Instead of passing the audio chunk individually through each effect, what we can
do instead is combine this exact chain into one function and let our compiler
backend of choice optimize it as if it were regular ol’ code.&lt;/p&gt;
&lt;p&gt;If we compare that to a generalized pipeline, things get a bit uglier
performance-wise.  You start to run for example into walking nodes dynamically,
and due to the amount of indirection in that, it can turn out rather slow.  You
may also need to check types and configuration at runtime as well, adding even
more overhead, and you may get branches that could literally never be taken
depending on what parameters are set.&lt;/p&gt;
&lt;p&gt;Specialization gets rid of a lot of that – it takes what is essentially runtime
knowledge and turns it into compile-time knowledge, where branches can
disappear, calls can be inlined, constants propagated, locality improved,
registers allocated more effectively, etc.&lt;/p&gt;
&lt;p&gt;You can take this idea even further, and have parameters that would normally be
some simple variables, be turned into constants, and the ones that are changing
dynamically to remain as variables.  Though depending on the program, if you are
going to have external influence over parameters such as scripts, for example,
this may not be such a good idea.&lt;/p&gt;
&lt;p&gt;This isn’t all sunshine and rainbows though, as there are some real downsides.
Those are in terms of invalidation and recompilation.&lt;/p&gt;
&lt;p&gt;You can have an entire effect chain but then you want to add a new effect; this
means that suddenly you have to recompile the whole thing again.  You may also
have multiple variants consuming even more memory.  So you need to be careful
depending on your requirements.  For example, if you want to make a DSP for a
microcontroller, it may not even be worth the effort implementing a code
generation system, and might as well put each effect behind a flag, and have the
specialization in the actual code.&lt;/p&gt;
&lt;p&gt;Ultimately, if you know about the pipeline ahead of time, consider turning that
runtime knowledge into compile-time knowledge.  This way you give the compiler
opportunities that a generalized architecture shrimply cannot provide.&lt;/p&gt;
</content:encoded><category>nerd</category><category>programming</category><category>kamastudio</category></item><item><title>Heron&apos;s approach to reflection</title><link>https://raung0.github.io/blog/herons-approach-to-reflection/</link><guid isPermaLink="true">https://raung0.github.io/blog/herons-approach-to-reflection/</guid><description>A look at how Heron handles code generation and reflection</description><pubDate>Tue, 01 Sep 2026 00:00:00 GMT</pubDate><content:encoded>&lt;p&gt;I have been working on a statically typed, memory-safe programming language
called Heron.  While going through many, many, &lt;em&gt;many&lt;/em&gt; stages of iteration of
the language’s design, I finally landed on a reflection system I am happy with.&lt;/p&gt;
&lt;p&gt;First though, I should set in stone what I think reflection is and what it
allows you to do.  Reflection, in my opinion, is a language feature that allows
the program to perform introspection as well as arbitrary code generation.  Note
that none of this relies on external tooling, such as Python scripts or other
programs like &lt;a href=&quot;https://doc.qt.io/qt-6/moc.html&quot;&gt;moc&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Reflection and how it should be done have been – like programmers normally love
to do – widely debated in how they should be implemented in programming
languages.  Some do it in a very primitive manner, bolting it onto the language
through many functions with different names, others create entire DSLs on top.
I wanted to avoid that and make the system as simple but as powerful as
possible.  Easy right? Well…&lt;/p&gt;
&lt;h1 id=&quot;the-approaches&quot;&gt;The Approaches&lt;/h1&gt;
&lt;p&gt;There are many approaches you could take while designing such a system.  Let’s
split the reflection into its two parts: the introspection and the code
generation.&lt;/p&gt;
&lt;p&gt;For introspection there are a couple of approaches that we can take:&lt;/p&gt;
&lt;h2 id=&quot;specialized-functions&quot;&gt;Specialized functions&lt;/h2&gt;
&lt;p&gt;This is the easiest but most verbose option – both in implementation and usage.
This approach relies on making various functions that only do one specialized
thing on some value representing the code.  Those functions are usually
implemented as compiler/interpreter built-ins, and as a result, you end up with
a huge compiler that contains a lot of boilerplate.&lt;/p&gt;
&lt;p&gt;For example, let’s say we want to iterate over all the fields of a structure and
print the names of its members. In C++26, that would look something like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-cpp&quot;&gt;struct GameState {
    float paddle_y[2];
    int button_state;
};

template&amp;lt;typename T&amp;gt;
void dump(T const &amp;amp;obj)
{
    template for (constexpr auto member :
        std::meta::nonstatic_data_members_of(
            ^^T,
            std::meta::access_context::current()))
    {
        std::cout &amp;lt;&amp;lt; std::meta::identifier_of(member) &amp;lt;&amp;lt; &apos;\n&apos;;
    }
}

int main() {
    GameState s { {2.0f, 4.0f}, 0b1101 };
    dump(s);
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We can already see the problem with this approach – it is extremely explicit!
This introduces a whole bunch of functions, like &lt;code&gt;nonstatic_data_members_of&lt;/code&gt; and
&lt;a href=&quot;https://cppreference.com/cpp/meta/reflection&quot;&gt;others&lt;/a&gt;.  This bloats up the meta
namespace, which makes finding the right function hard, especially in
autocomplete (look at the &lt;code&gt;is_&lt;/code&gt; functions!), and the function names are also
very explicit making the code harder to parse.&lt;/p&gt;
&lt;h2 id=&quot;macros&quot;&gt;Macros&lt;/h2&gt;
&lt;p&gt;Macros are able to help with this issue, but are much more limited.  The usual
approach is to create two files: one with the actual code, and one with the
member definitions. So, you end up with something like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-cpp&quot;&gt;// members.def
X(float, paddle_y, [2];)
X(int, button_state, ;)

// main.cpp
#define X(type, name, type_post) type name type_post
struct GameState {
#include &quot;members.def&quot;
};
#undef X

#define X(type, name, type_post) std::cout &amp;lt;&amp;lt; #name &amp;lt;&amp;lt; &apos;\n&apos;;
void dump(GameState const &amp;amp;t) {
#include &quot;members.def&quot;
}
#undef X

int main() {
    GameState s { {2.0f, 4.0f}, 0b1101 };
    dump(s);
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;We can already see the first limitation of this approach.  Firstly, &lt;code&gt;dump&lt;/code&gt; is no
longer templated, so it will only work with &lt;code&gt;GameState&lt;/code&gt;.  If we wish to make
&lt;code&gt;dump&lt;/code&gt; work with other types, we need to create overloads, which leads to code
duplication.  Secondly, all our member definitions &lt;em&gt;have&lt;/em&gt; to live in another
file if we do not want to repeat ourselves.  Not only that, we also have to
define the actual code to be generated.&lt;/p&gt;
&lt;h2 id=&quot;ast-data-structure&quot;&gt;AST data structure&lt;/h2&gt;
&lt;p&gt;All of this leads to my choice for introspection – an AST data structure in the
standard library.  This may seem a bit of a weird choice at first, since you are
exposing compiler internals, right?&lt;/p&gt;
&lt;p&gt;Nope!  The AST can be completely independent from the compiler’s representation.
You can define a generic, detailed AST that covers pretty much every primary
aspect of the language and have the compiler transform its own representation to
it.  This is basically what I went for in my language.  So, the code turns out
something like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;GameState :: struct {
  paddle_y: [2]f32;
  button_state: int
};

dump :: fn(ty: type) {
  tree := @get_ast(ty);
  match declaration in tree {
    case rf.Declaration -&amp;gt; match type_decl in declaration.children[0] {
      case rf.TypeDeclaration -&amp;gt; match fields in type_decl.children[0] {
        case rf.Block -&amp;gt; for child in fields.children {
          match field in child {
            case rf.Declaration -&amp;gt; if !field.comptime do {
              fmt.println(&quot;{}&quot;, field.name);
            }
          }
        }
      }
    }
  }
};

main :: fn do dump(GameState);
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Notice that tree is just an AST node.  This means we can now simply get its
declarations and print them normally.  Yes it’s a bit more explicit than I would
like it to be, but the language is still WIP, so there are some missing things
(like we shouldn’t require a whole match block just for one case!).  But you can
see how, by using existing language features, this code becomes much easier to
parse, without having to bloat up the namespace, as you only get what you need
because of the AST node fields.  This could probably be even more simplified if
I make the AST node members not point back to the &lt;code&gt;AST&lt;/code&gt; where applicable, but
instead dedicated types.&lt;/p&gt;
&lt;h1 id=&quot;code-generation&quot;&gt;Code generation&lt;/h1&gt;
&lt;p&gt;OK, so we got introspection, what about generating code?  We already saw a bit
of that when doing introspection with macros, but there are a few other options
as well:&lt;/p&gt;
&lt;h2 id=&quot;specialized-functions-1&quot;&gt;Specialized functions&lt;/h2&gt;
&lt;p&gt;The first obvious one is, again, specialized functions.  Have a function to
generate whatever.  This has the same issues as with introspection, lots of
boilerplate, and hard and clunky to use.  This is what I initially thought for
heron, have a bunch of functions or maybe a builder-style API to generate code.
When it became time to actually think about reflection properly, I realised the
immense effort of implementing this approach, so I didn’t take it.&lt;/p&gt;
&lt;h2 id=&quot;token-streams&quot;&gt;Token streams&lt;/h2&gt;
&lt;p&gt;One of the first places I saw this approach was in HolyC of all places.
Basically, the language has a special compiler directive that’s called &lt;code&gt;#exe&lt;/code&gt;,
which runs code at compilation time.  This can be used to insert into the token
stream of the compiler strings and achieve code generation that way.  For
example, you can generate a whole bunch of definitions like this:&lt;/p&gt;
&lt;pre&gt;&lt;code class=&quot;language-holyc&quot;&gt;#exe {
	I64 i = 0;
	for (i = 0; i &amp;lt; 10; i++)
		StreamPrint(&quot;#define VALUE_%d (%d)\n&quot;, i, i++);
};
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;While this is quite powerful, it can also introduce a bunch of mistakes.  What
happens if you create a syntax error by accident?  Now you have to go, find the
code that generates the string and fix it (if you can even find the source of
the bad part of the string, that is).  This, in my opinion, in the long-term,
leads to bad ergonomics.&lt;/p&gt;
&lt;h2 id=&quot;getting-lispy-with-quotes&quot;&gt;Getting lispy with quotes&lt;/h2&gt;
&lt;p&gt;Lisp, however, has a very nice approach – quotes.  For example, you can have a
little function that generates a bit of quote to increase the value of a
variable:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;(defmacro inc (place)
  `(setq ,place (+ ,place 1)))

(inc some_counter)
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here you can see what it does, the &lt;code&gt;,id&lt;/code&gt; basically inserts whatever that value
is literally in place, so what that call to &lt;code&gt;inc&lt;/code&gt; becomes is essentially this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;(setq some_counter (+ some_counter 1))
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;A similar approach can be taken with other programming languages. For example in
heron, if we want to have a similar &lt;code&gt;inc&lt;/code&gt; function:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;inc :: fn(place: code) -&amp;gt; code do &apos;{
  ,{place} = ,{place} + 1
};

main := fn {
  mut some_counter := 0;
  @insert(inc(&apos;{ some_counter }));
};
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Notice that we have a function that takes in code and returns other code. Code
is a first class citizen in Heron, and is treated almost like every other type.&lt;/p&gt;
&lt;p&gt;My language also introduces &lt;code&gt;,@{}&lt;/code&gt; and &lt;code&gt;@emit&lt;/code&gt;, which can be used to run code
that emits other code. For example:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;use core.fmt;

make_prints :: fn -&amp;gt; code {
  return &apos;{
     ,@{
        for i in 1..&amp;lt;4 {
           @emit(&apos;{
              fmt.println(&quot;{}&quot;, ,{i});
           })
        }
     }
  }
};

main := fn {
  @insert(make_prints())
};
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Which essentially makes &lt;code&gt;main&lt;/code&gt; turn into:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;main := fn {
  fmt.println(&quot;{}&quot;, 1);
  fmt.println(&quot;{}&quot;, 2);
  fmt.println(&quot;{}&quot;, 3);
};
&lt;/code&gt;&lt;/pre&gt;
&lt;h1 id=&quot;conclusion&quot;&gt;Conclusion&lt;/h1&gt;
&lt;p&gt;All in all, the approach I settled on is surprisingly simple: use a regular AST
data structure for introspection and make code a first-class citizen for
generation.&lt;/p&gt;
&lt;p&gt;I especially like that with this approach, &lt;em&gt;&lt;strong&gt;everything doesn’t feel like it’s a
whole separate thing embedded into Heron and that it’s actually part of the
language – which it is&lt;/strong&gt;&lt;/em&gt;.  There’s no huge list of compiler intrinsics to learn,
no external preprocessing to be done and none of it is stringly typed.  Most of
this is already expressed with features Heron already has, things like structs,
typed unions, functions and basic control flow.  While the compiler does have to
provide &lt;em&gt;some&lt;/em&gt; intrinsics like &lt;code&gt;@get_ast&lt;/code&gt; or &lt;code&gt;@emit&lt;/code&gt;, they are far fewer in
number than an approach based on specialized functions.&lt;/p&gt;
&lt;p&gt;There are still lots of quirks to figure out, better ergonomics, how diagnostics
should work so that they are nice and readable, but this design remains, in my
opinion, a very, &lt;em&gt;very&lt;/em&gt; solid foundation for what’s to come in future revisions
of the language.  After all of this, I will probably be able to just stop
redesigning the reflection system from scratch for months.&lt;/p&gt;
&lt;p&gt;Probably.&lt;/p&gt;
</content:encoded><category>nerd</category><category>programming</category><category>heron</category><category>compilers</category></item></channel></rss>