Semurg - a CPU-native data engine built with Elixir and Rust

Hi All,

Long-time reader, finally made an account.

I wanted to introduce something my team and I have been building, mostly because the core of it is an Elixir and Rust bet and I would really like this community’s take on it.

Semurg holds every type of data (relational, graph, object, document, search, vector, time-series) as one fixed 64-byte binary container, in a single engine over one copy of the data, CPU-only, on hardware you control. The engine is a commercial, closed-source product, but the installer and a full benchmark suite are now open, and the first node is free forever. It is live with no login at one.semurg.io, and you can now stand up your own node in one command and benchmark it yourself:

Install a node (prebuilt engine, no compiler, no Rust, no separate Erlang): GitHub - One-Semurg/Semurg-Install: Official one-command installer for Semurg. Downloads and runs the released Semurg engine — your first node is always free. Install script and release binaries only; this repo does not contain the engine source. · GitHub

Benchmark it on your own box against the engines you already run, equal-answer gated and bit-exact: GitHub - One-Semurg/Semurg-Benchmark-Suite: Run-it-yourself benchmarks comparing Semurg to leading open-source databases across eleven data domains - on your own hardware, your own results (no published numbers). · GitHub

The part that belongs on this forum is how it is put together:

BEAM conducts, Rust does the heavy lifting. Elixir and OTP own supervision, scheduling and fault tolerance. The per-byte work runs in Rust NIFs, one concurrent arm per physical core, with no shared write lock between them. We hand each NIF a whole batch and fan out inside a single crossing rather than paying the boundary cost per item, and only tokens and offsets ever cross, never materialised data. Long jobs go on dirty schedulers, and every NIF has a pure-BEAM fallback, so a missing or failed native artifact costs us throughput rather than taking the VM down.

The 64 bytes are not arbitrary. One container is a CPU cache line, 64 of them is a 4KB disk to RAM page, so the storage format is aligned to what the hardware actually moves.

Agents are just supervised BEAM processes over the same store, so a swarm gets isolation and concurrency from the runtime rather than from containers.

On scope, so I am not overselling it: the substrate and the gateway are live today, and you can run a node yourself now. With local models your data never leaves your box, so tokenising PII before an external model call, which I used to do, is not something I need any more, and I have dropped it. Growing CPU-native models on the same substrate is still research, and I try to keep that line clear.

On numbers, rather than ask you to trust a chart: the benchmark suite stands up each open-source incumbent from its official image, on the same deterministic data, and only counts a result if its answer hash matches an independent reference. It prints the losses straight next to the wins. The OLAP lane shows Semurg’s raw-scan loss to DuckDB honestly, next to the O(1) fold win. The graph lane has an out-of-core regime where the in-RAM engines DNF on a fixed small memory budget and Semurg finishes the traversal at flat memory, which is the property I actually care about.

Happy to go into as much depth as anyone wants on the orchestration and the BEAM side. The storage layout and IO mechanics are the bit we keep closed, so I will be vaguer there, but everything above that is fair game.

What I would most like is to hear from anyone who has run BEAM and Rust NIFs in production, especially dirty schedulers and large batched crossings. What caught you out? I have a feeling we are not the first to learn some of these the hard way.

Thanks

5 Likes

Quick update for anyone who saw this back in July, since a few things changed.

The site moved. semurg.io is retired. Everything is now at one.semurg.io.

Anyone can run it themselves now. Back then it was live to poke at, but closed and hosted by me. That changed. The engine is still closed source, but the installer and a full benchmark suite are now public, and the first node is free forever. You stand up a real node on your own hardware in one command, and you benchmark it against whatever you already run, with no account, no sales call, and nobody in the loop but you.

Install a node (prebuilt engine, no compiler, no Rust, no separate Erlang to set up):

Benchmark it on your own box (equal-answer gated, bit-exact, no fake numbers; it stands up each open-source incumbent from its official image and runs the same queries on the same deterministic data, and prints the losses straight next to the wins):

Here is the actual reason I am posting, though. I want your feedback to make the engine better. It is early, and the fastest way I know to uplift it is to put it in front of people who will run it on workloads I have not thought of and tell me exactly where it falls over. If you clone the suite and Semurg loses a lane on your box, that is the single most useful thing you can send me, because every one of those becomes the next thing I fix.

A couple of honest places it will lose today, so you know I mean it. The OLAP lane prints Semurg’s raw-scan loss to DuckDB straight, right next to the O(1) fold win, because a columnar warehouse is built to win that scan and I would rather show it than hide it. Where it does something the others cannot is the out-of-core graph regime: on a fixed small memory budget the in-RAM engines DNF and Semurg finishes the traversal at flat memory, which is the property I actually care about. Run it, break it, and tell me what caught you out or where it lost.

On the BEAM and Rust side from the original post: it is the same engine, still one concurrent NIF arm per physical core, dirty schedulers for the long jobs, batched crossings with only ids and offsets across the boundary, and a pure-BEAM fallback per NIF. The single-node install is just that engine packaged so you can run it without me in the loop, and watching where it breaks in other people’s hands is how it gets good.

Final updates should be pushed to the engine and the benchmark suite in days and then is ready for full testing and benchmarking.