Hi,
I am building BeamTrail, a small BEAM library for durable business processes: retries, timers, signals, approval deadlines, operator controls, and crash recovery, backed by PostgreSQL.
It is not a Temporal clone and not a job queue. The goal is an embedded runtime for Erlang/Elixir/Phoenix apps that already run PostgreSQL and need long-running workflows without operating a separate workflow service.
The basic idea is:
- Active runs are supervised OTP processes.
- Business progress is stored as an append-only PostgreSQL event stream.
- State is rebuilt by reducing events.
- Snapshots and indexed projections are optimizations only.
- Leases, fencing tokens, and expected sequence checks protect concurrent writers.
- A scanner recovers unfinished runs whose lease is missing or expired.
The current use cases I am aiming at are things like:
- order fulfillment with retryable external calls;
- payment or webhook-driven sagas;
- approval flows with deadlines;
- long-running business processes that must survive VM restart;
- AI agent/tool runs where the session memory is separate, but tool calls, approvals, timers, and recovery need durable
execution boundaries.
Current implemented pieces:
- PostgreSQL durable adapter and memory test adapter.
- Append-only event log and reducer.
- Expected-sequence append checks.
- Per-run PostgreSQL append lock.
- Lease renewal and fencing tokens.
- Supervised per-run
gen_statemactive runner. - Scanner recovery with indexed recoverable-run projections.
- Retries, timeouts, and crash-atomic failure decisions.
- Cancel, park/resume, and manual requeue.
- Version mismatch gates for in-flight attempt replay and decider workflows.
- Durable signals and scanner-driven durable timers.
- Executable approval-deadline pattern.
- Crash recovery and PostgreSQL stress examples.
- EUnit, PostgreSQL integration tests, xref, Dialyzer, and secret scan in CI.
Current limits:
- No DAG, fan-out/fan-in, child workflows, or parallel command batches yet.
- No first-class human task assignment UI.
- No HTTP API or browser UI.
- No SQL-native JSON payload inspection.
- No built-in external side-effect deduplication.
- No exactly-once execution of user callbacks.
The at-least-once boundary is explicit. BeamTrail provides stable idempotency keys, but workflow code must use them when calling external systems. If a VM dies after a callback performs an external side effect but before BeamTrail records the outcome, the same attempt can be re-entered with the same idempotency key.
How I currently think about the positioning:
- If you only need background jobs, Oban is probably the better tool.
- If you need a mature multi-language workflow platform, Temporal is probably the better tool.
- BeamTrail is trying to sit in the smaller space between them: an embedded BEAM runtime for durable business processes, backed by PostgreSQL and shaped around OTP supervision.
What I would like feedback on:
- For Phoenix/Elixir apps, does this fill a real gap between Oban jobs and Temporal?
- Are the current APIs and semantics understandable from an Elixir user’s point of view, even though the library is written in Erlang?
- Would durable signals, timers, approval deadlines, and crash recovery be enough to try this in a small internal workflow?
- What would be the minimum missing feature before you would consider experimenting with it?
- Is PostgreSQL as the durable boundary acceptable, or would you expect a different storage/runtime shape?
Useful docs:
Repository:
There is also a GitHub Discussion for longer-form follow-up:
I am especially interested in practical feedback from people who run Phoenix/Elixir apps with background jobs, workflows, approvals, webhooks, or long-running business processes.
Thanks.






















