StackMap
Subscribe
Explore / lossless-memory
aru-labs

lossless-memory

Lossless long-term memory for a personal AI: verbatim JSONL logs, time-first search over SQLite FTS5 (sqlite-vec as last resort) and a tiny 'where are we now' index injected every turn.

143 13 Python MITupdated 18 days ago
View on GitHubDispute this mapping →
Curator's take

Choose it for one person, one assistant, one machine, when 'what exactly did we say last Tuesday' matters more than distilled facts — it never summarizes, and time phrases narrow the range before anything is ranked. Not for multi-user products, agent fleets or benchmark-chasing: no server, no published evals. If you want extracted facts, profiles and forgetting, supermemory is the opposite philosophy.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside lossless-memory. Ranked by curator confidence.

alternativealternativealternativeMemMoltdeja-vusupermemorylossless-memory
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md3 min read

lossless-memory

Lossless long-term memory for a personal AI — never summarize, keep every line, and put a timestamp on everything.

Most long-term memory systems for AI do one of two things: they summarize conversations into compact notes, or they embed them and retrieve "similar" chunks. Both lose the thing that matters most to a person who talks to the same AI every day: what was actually said, and when.

This project takes the opposite position.

  • Keep every line. Raw conversation logs are stored in full. Nothing is summarized, ever. Summaries are a map; the log is the territory.
  • Timestamp everything. Every record — utterance, action, document chunk — carries a timestamp, and every index is built on top of that time axis. We call this the Temporal Backbone.
  • Search by time first, words second. "Yesterday evening, about the budget" is a valid query. The time phrase narrows the range; the words rank within it. Results come back in chronological order, unsummarized, with their timestamps.
  • Inject "where we are" every turn. A small index called LLL tells the model which topic the conversation is in right now, so identity and context survive context-window compaction and session boundaries.

The design lineage goes back to December 2025 — the first ancestor of this system (a memory-inheritance tool for an earlier AI) ran that month, and a predecessor system carried the same ideas in daily use from January 2026. This implementation has been running every day since July 2026 for a single user, as the memory of one AI assistant, with raw logs reaching back to June 2026. It is small, boring, and it works. The failures along the way are documented too — see docs/lessons.md.


What this is / what it is not

It is:

  • A local, file-based long-term memory layer: JSONL logs + SQLite (FTS5 for exact search, sqlite-vec for semantic search).
  • A single query entry point that understands time expressions and restricts the search range before ranking.
  • A "current position" index (LLL) designed to be injected into the model's context on every turn.
  • Designed for one person and one AI, running on one machine. No server, no cloud.

It is not:

  • A vector database wrapper. Semantic search is the last resort here, not the first.
  • A summarizer. There is deliberately no summarization step anywhere in the pipeline.
  • A benchmark-driven research system. There are no published benchmarks. What is here is a working implementation and its operating record.

The three pillars

1. Lossless raw log

Every conversation turn is converted into a fixed seven-field record and appended to a per-day JSONL file:

ts        ISO-8601 timestamp (UTC)
actor     who spoke (configurable names)
role      user | assistant | system
type      text | action | meta
text      the content, verbatim
model     model identifier, if known
session   session identifier

The raw logs are the source of truth. Every index below can be deleted and rebuilt from them. Nothing else is required to survive.

2. Temporal Backbone

Time is not metadata here; it is the primary axis.

  • The exact-match index (SQLite FTS5, bigram tokenized for Japanese and English) stores the timestamp alongside every row.
  • The query parser understands time phrases — relative ones such as yesterday, last week, 3 days ago, in July, this morning (in English and Japanese), and absolute dates such as 2026-07-19 (any language) — and converts them into a range before any ranking happens. "Yesterday" means yesterday in your timezone (timezone in config.json).
  • If a time phrase is present, results are restricted to that range and returned in chronological order. Semantic search is only used when the exact index returns too little inside the range, and the fallback is reported honestly in the output header.

The practical effect: the AI can answer "what did we decide last Tuesday night?" with the actual l