A database with an algebraic type system and first-class schema migrations. https://tempest.jofh.me
  • Rust 92%
  • Vue 4%
  • TypeScript 3.3%
  • SCSS 0.6%
Find a file
2026-06-01 20:47:47 +02:00
docs/scenarios docs: add scenario for simple posts schema 2026-06-01 20:47:47 +02:00
site docs(site): correct statement on primary key uniqueness 2026-04-30 18:08:54 +02:00
tempest_core feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
tempest_db feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
tempest_engine chore: delete most of tempest_engine to start fresh 2026-05-26 14:08:33 +02:00
tempest_io feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
tempest_kv feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
tempest_repl feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
tempest_rt fix: issue with certain runtime fs ops on uncompleted polls 2026-05-22 16:23:11 +02:00
tempest_tql feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
tempest_wasm feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
.gitignore feat: wasm io backend, vue+nuxt page, wasm worker based browser repl 2026-04-11 15:54:26 +02:00
.rumdl.toml docs: rewrite README.md, use rumdl as md linter 2026-04-14 01:26:47 +02:00
build.sh feat: wasm io backend, vue+nuxt page, wasm worker based browser repl 2026-04-11 15:54:26 +02:00
bun.lock feat: wasm io backend, vue+nuxt page, wasm worker based browser repl 2026-04-11 15:54:26 +02:00
Cargo.lock feat: ansi tables for bold headers 2026-04-11 17:04:40 +02:00
Cargo.toml feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00
LICENSE chore: unify workspace metadata and license to FSL-1.1-ALv2 2026-04-02 02:41:47 +02:00
package.json feat: wasm io backend, vue+nuxt page, wasm worker based browser repl 2026-04-11 15:54:26 +02:00
README.md feat: make ready for cargo publish 2026-04-23 20:50:36 +02:00

TempestDB

wakatime

Important

This is a learning project under active development. It is not possible to use this for anything in production for the time being.

This document also lays out a few features that are either under development or may not even exist.

Tempest is a database I am building in order to modernize the database experience, by approaching things differently:


Installation

Install the CLI with Cargo:

cargo install tempest-db

The binary is available as tempest after installation.

Documentation

The demo and full documentation are available at tempest.jofh.me.


The Type System

The main thing that I've found to make me unhappy with current SQL database systems:

  • They make everything nullable by default (which often results in bugs if left unchecked in queries)
  • They don't have support for proper enums (ones that some may refer to as tagged unions)

I would go so far as to say that nullability will not be fixed by making everything NOT NULL by default, but rather by getting rid of the concept that that is NULL, and instead replacing it with an Option[T] enum, similar to how Rust does it.

This will also get rid of SQL's three-valued logic system and the problems that come with it. In case you did not know this, SQL actually has three values of truthiness, which are TRUE, FALSE, and UNKNOWN. Comparing NULL with NULL is not actually TRUE, but UNKNOWN instead. It makes sense in a way, because if you compare the emails of two users, and both are NULL, you don't know if they have the same email; it does not make sense mathematically and I would consider it dangerous behavior. I consider handling absence explicitly to be superior for correctness in all applications that are developed.

What this also means, is that we'll require generics to make the T part of the Option work. On that matter, I was also thinking about adding generic bounds, so that we can have something like a Comparable interface/trait that allows us to specify functions like min[T] or max[T].

Note

I will probably go with monomorphization of all generic variants, since a dynamic dispatching features to support things like heterogeneous lists is not something reasonable to expect in a database.

Data Shape and Data Storage

A type defines the shape of data, like structs and enums. A table attaches storage to a type. This means types can be embedded, reused, and referenced independently of any particular table, making user-defined functions that take or return table entries directly very easy to define.

create type Address struct {
    street  : String,
    city    : String,
    country : String,
}

create type User struct {
    id       : Int64,
    username : String,
    email    : String,
    age      : Int8?,     // `?` suffix = `Optional` - may be `Some(value)` or `None`
    address  : Address,   // embedded type by default, which would make join not required
}

// `table` attaches storage to a type
create table users : User {
    // We declare the primary key in the table, not the type, because it is a concern
    // of the storage strategy and has nothing to do with the shape of the data
    primary key (id),
    // This will make the field "address" be a reference to the table "addresses"
    reference address to addresses,
}

The rule is making types PascalCase, keywords lowercase, and : means "has type."

This query language is what I've been calling the Typed Query Language or TQL for short - evidently named after SQL, with a focus on the improved type ergonomics.

Syntax Features for Option[T] in TQL

If a field can be absent, the type says so, and the query language requires you to handle it.

// Pattern-match style: filter and bind in one step
select id, username
from   users
where  age? > 18;
//       ^ `?` in a query unwraps the Optional,
//         implicitly filtering out rows where age is None

For cases where you want to handle None explicitly:

select id, username, age
from   users
where  age is Some(a) and a > 18;
//              ^ binds the inner value to `a` only when present

Migrations

The idea that I'm still exploring is making schemas declarative instead of imperative. This would mean, you write something like declare type instead of create type and you would then let TempestDB's tooling handle the migration for you, similar to what ORMs provide you with, but out of the box and without the additional layer of abstraction on top.

Performance and Safety

TempestDB is built with peformance and safety in mind from the start, which obviously starts with choosing Rust as the implementation language. It does not automatically make everything fast, if your implementation is bad though.

I've decided to build the relational database on top of a flat key-value store. The KV layer is implemented as a log-structured merge tree, allowing for write-throughput beyond what is possible with regular B-Trees, because it avoid the frequent disk-seeks that come with random-access and instead results in sequential writes, which is faster on modern solid-state drives.

Optimizing I/O Throughput with io_uring

The Linux-specific io_uring API allows us to efficiently do asynchronous disk and network I/O without any syscalls that would block execution. Instead, it uses two ring buffers, one for I/O submissions, and one for completions, that sit in userspace, allowing the kernel as well as us to access it simultaneously.

The kernel continuously pulls submission queue entries (SQEs) from the submission queue (SQ), processes them in the background, and when done, write them into the completion queue (CQ) in the form of completion queue entries (CQEs).

This model allows us to just write SQEs in there and check back every now and then if an I/O operation has been completed by now, while also letting us submit others. With regular I/O operations, our thread would be sleeping until completion, which I consider to be absolutely unacceptible for a modern database.

This brings one problem with it though: io_uring is instanced once per thread. You cannot just switch the thread context and read from another threads CQ. This brings up a fundamental conflict with the async runtime Tokio, because when you've enabled the automatic multi-threading feature of the runtime, it will spawn a certain number of worker threads and cooperatively schedules the background tasks, which may end up being executed on an entirely different thread from where they were started at. This is normally fine, but in the case of io_uring, when originally trying to integrate it with tokio-uring, I've noticed that it did not work as simple as I thought it would. It requires us to spawn a special tokio-uring worker thread, in which we do all of the I/O.

My first instinct was using channels to communicate I/O requests with that worker thread, but it felt awkward and is also fairly slow, because we move a lot of data around all the time. Early prototypes also used Arc and Mutex everywhere to synchronize the data structures across thread boundaries. This meant that we did not even do any proper multi-threaded compute, since the threads where usually just waiting in line anyways.

So I went a different route and rethought the whole architecture. It kind of forced me to move towards what is called a shared-nothing architecture - each database shard owns its data and communicates via message passing, no locks required. It turned out to be both easier to reason about and more performant, which is the same bet ScyllaDB, TigerBeetle, and Redpanda made and won against their locked counterparts.

As of right now I only have a single shard. Multi-shard is not a strict requirement for a minimum viable product, which is why I've obviously deferred that for now, since distributed systems introduce a whole separate list of questions to answer. A single shard also runs fairly fast, considering that for databases, I/O is the main bottleneck, which we're already doing asynchronously, so this single shard is probably more performant than a few threads with locked datastructures and synchronous I/O would be.

Building a Custom I/O-agnostic and Deterministic Async Runtime

-- TODO --