Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ edition = "2024"
authors = ["Boog900"]
license = "MIT"
repository = "https://github.com/Cuprate/tapes"
description = "A minimal database for storing data in contigious tapes"
description = "A minimal database for storing data in contiguous tapes"

[dependencies]
slab = { version = "0.4" }
Expand Down
53 changes: 41 additions & 12 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,18 +2,47 @@

A specialised database for storing data in contiguous tapes.

Each environment supports multiple independent tapes, with ACID updates across them. There are 2 types of tapes: fixed-sized and
blob. Fixed-sized tapes store fixed sized values, which allows arbitrary lookup of values by their index, blob tapes however is
just a contiguous slice of bytes so to access data you must keep its index.
A tape is an append only log of data. This database is not a generic key/value store, it is
specialised for data that builds on top of previous data, in a way that old data is only removed
if the data created after it is removed too, and removing data is less common than adding it.
An example of such data would be a typical blockchain.

## Supported Operations
Each [`Tapes`] instance supports multiple independent tapes, with ACID updates across them.

This database is not a generic key/value store. This crate is highly specialised for storing data that builds on top of previous
data, in a way that old data will only be removed if data created after it is also removed and that removing data is less
common than adding data. An example of such data would be an average blockchain.
## Tapes

Instead of there being a reader and writer transactions, write transactions are split into append and pop. This means there
are 3 transaction types: read, append and pop. Splitting the writer transaction like this makes the database more efficient
at the cost of not being able to do a single atomic rewrite of data. It is important to note though that removing data and
then writing more is still ACID, it's just the database could be left with the data being removed without the new data being
written.
There are 2 kinds of tapes: fixed sized tapes store fixed sized values, which allows lookup of
values by their index, blob tapes are a contiguous slice of bytes, so to access data you must
keep its index.

A [`WholeBlobTape`] stores all its data in a single file. A [`RollingBlobTape`] stores its data
in multiple files of at most a configured size, and deletes the files that hold removed data
once no reader needs them anymore. On a whole tape removing data from the front does not free up
disk space, on a rolling tape it does. Use a rolling tape when the tape grows without bound and
you need to remove old data FIFO, if not use a whole tape.

A [`FixedSizedTape`] is a handle to a blob tape that reads and writes fixed sized entries instead
of raw bytes, the entry type must be plain data with no padding [`bytemuck::Pod`].

A [`CachedBlobTape`] is a wrapper around a blob tape that keeps the top of the tape in memory,
this speeds up access to recent data and reduces disk I/O. These are both wrappers, so they can
be combined, for example a fixed sized tape over a cached rolling tape.

You probably want to always use a [`CachedBlobTape`] to prevent doing slow direct I/O.

## Transactions

Write transactions are split into append and pop, this means there are 3 transaction types:
read, append and pop. Splitting the writer transaction like this makes the database more
efficient at the cost of not being able to do a single atomic rewrite of data. Removing data and
then writing more is still ACID, it's just the database could be left with the data being
removed without the new data being written.

## Persistence

Commits are persisted according to the [`Persistence`] mode they are committed with:

- `Buffer` writes to the OS buffer only, it is not durable, data committed with it can be lost
on a crash until it is flushed to disk by a later commit with `SyncData` or `SyncAll`.
- `SyncData` syncs the file contents to disk.
- `SyncAll` syncs the file contents and file metadata to disk.
61 changes: 61 additions & 0 deletions src/io_helpers.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
use std::{fs::File, io};

pub(crate) fn read_exact_at_file(file: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;

file.read_exact_at(buf, offset)
}

#[cfg(windows)]
{
use std::os::windows::fs::FileExt;

let mut buf = buf;
let mut offset = offset;
while !buf.is_empty() {
match file.seek_read(buf, offset) {
Ok(0) => {
break;
}
Ok(n) => {
buf = &mut buf[n..];
offset += n as u64;
}
Err(e) => {
return Err(e);
}
}
}

if !buf.is_empty() {
Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"failed to fill the whole buffer",
))
} else {
Ok(())
}
}
}

pub(crate) fn write_all_at(file: &File, buf: &[u8], offset: u64) -> io::Result<()> {
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;

file.write_all_at(buf, offset)
}
#[cfg(windows)]
{
use std::os::windows::fs::FileExt;

let n = file.seek_write(buf, offset)?;
if n != buf.len() {
return Err(io::Error::other("Failed to write all bytes to tape"));
}

Ok(())
}
}
18 changes: 14 additions & 4 deletions src/lib.rs
Original file line number Diff line number Diff line change
@@ -1,18 +1,28 @@
#![doc = include_str!("../README.md")]

mod metadata;
mod ring_buffer;

mod io_helpers;
mod tapes;
mod traits;

pub use tapes::{
BlobTape, FixedSizedTape, TapeOpenOptions, Tapes, TapesAppendTransaction, TapesReadTransaction,
TapesTruncateTransaction,
CachedBlobTape, CachedTapeOpenOptions, FixedSizedTape, RollingBlobTape, RollingTapeOpenOptions,
Tapes, TapesAppendTransaction, TapesReadTransaction, TapesTruncateTransaction, WholeBlobTape,
WholeTapeOpenOptions,
};
pub use traits::{TapesAppend, TapesRead, TapesTruncate};
pub use traits::{BlobTape, TapesAppend, TapesRead, TapesTruncate};

/// How a commit is persisted.
///
/// `Buffer` is not durable, data committed with it can be lost on a crash until it is flushed to disk
/// by a later commit with `SyncData` or `SyncAll`.
#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)]
pub enum Persistence {
/// Writes to the OS buffer only, not durable.
Buffer,
/// Syncs the file contents to disk.
SyncData,
/// Syncs the file contents and file metadata to disk.
SyncAll,
}
23 changes: 22 additions & 1 deletion src/metadata.rs
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,19 @@ use slab::Slab;

use crate::Persistence;

pub type ActiveMetadata = HashMap<Box<str>, u64>;
pub type ActiveMetadata = HashMap<Box<str>, TapeMetadata>;

#[derive(Clone, Copy, Default, BorshDeserialize, BorshSerialize)]
pub struct TapeMetadata {
pub len: u64,
pub start: u64,
}

pub struct MetadataGuard {
active_metadata: Arc<ActiveMetadata>,
metadata: Arc<Metadata>,
reader_slot: usize,
pub epoch: u64,
}

impl Deref for MetadataGuard {
Expand Down Expand Up @@ -113,6 +120,7 @@ impl Metadata {
active_metadata: inner.active_metadata.clone(),
metadata: Arc::clone(self),
reader_slot,
epoch,
}
}

Expand All @@ -134,6 +142,16 @@ impl Metadata {

Ok(())
}

pub fn oldest_reader_excluding_reader(&self, guard: &MetadataGuard) -> Option<u64> {
self.inner
.lock()
.reader_epochs
.iter()
.filter(|(r, _)| *r != guard.reader_slot)
.map(|(_, r)| *r)
.min()
}
}

struct MetadataBackingFiles {
Expand Down Expand Up @@ -213,6 +231,7 @@ impl MetadataBackingFiles {

#[derive(BorshSerialize, BorshDeserialize)]
struct StoredMetadata {
version: u32,
hash: [u8; 32],
epoch: u64,
tapes: Vec<u8>,
Expand All @@ -221,6 +240,7 @@ struct StoredMetadata {
impl Default for StoredMetadata {
fn default() -> Self {
StoredMetadata {
version: 1,
hash: [0; 32],
epoch: 0,
tapes: borsh::to_vec(&HashMap::<Box<str>, u64>::new()).unwrap(),
Expand All @@ -244,6 +264,7 @@ fn serialise_metadata(epoch: u64, metadata: &ActiveMetadata) -> Vec<u8> {
let hash = hasher.finalize().into();

borsh::to_vec(&StoredMetadata {
version: 1,
hash,
epoch,
tapes: tapes_bytes,
Expand Down
Loading
Loading