---
title: "The SPDF 4.1 format, a library you can take with you"
description: "An .spdf is a gzip-compressed SQLite database holding the document already read: text, anchors, sections, figures, vectors and provenance. It opens without Scholaris and offline."
url: https://scholaris.joseluissaorin.com/en/knowledge/spdf
markdown: https://scholaris.joseluissaorin.com/en/knowledge/spdf.md
lang: en
alternate_es: https://scholaris.joseluissaorin.com/saber/spdf.md
updated: 2026-10-07
author: José Luis Saorín Ferrer (https://joseluissaorin.com)
---

# The SPDF 4.1 format, a library you can take with you

> An .spdf is a gzip-compressed SQLite database holding the document already read: text, anchors, sections, figures, vectors and provenance. It opens without Scholaris and offline.

## What an .spdf is

An **.spdf** file (Scholaris PDF) stores a document already read: not just the original but everything Scholaris got out of it, so it can be searched and cited elsewhere without reading it again or paying for it twice.

Inside it is a **gzip-compressed SQLite database** (not a ZIP). Any language with SQLite opens it; uncompressed files are accepted too. The current version is **4.1**; the version is stored in the `spdf` table under the key `spdf_version`. Old SPDFs (versions 1 to 3, from the first Scholaris) are migrated on the fly when opened.

Each exported .spdf holds one document. The same schema is what each user's library uses on the server.

## The tables

| Table | What it stores |
| --- | --- |
| spdf | Key and value: version, creation date, generator, hash of the original |
| documentos | Type, record (JSON), SHA-256 hash, MIME type, size, number of units, duration, title, authors, year and language |
| unidades | The citable units (pages, audio stretches, slides): anchor, text, notes, header and footer, image, confidence, printed page, times and per-word timings |
| secciones | The section tree |
| fragmentos | The passages that are searched and cited: text, context, section, start and end anchors, and the modernised-spelling layer (texto_busqueda) |
| fragmentos_fts | FTS5 full-text index over text, context, section and modernised spelling, accent-insensitive |
| figuras | Figures, plates and video frames, with region, caption and description |
| espacios | Vector spaces: provider, model, version, dimensions, normalisation and modalities |
| vectores | One vector per target (fragment, unit or figure) and space, as little-endian float32 |
| blobs | The original and the images (only in exported files) |
| procedencia | The log of how it was read: each phase, which provider did it, how long it took and when |

Version 4.1 added the `texto_busqueda` column: a shadow of the text in modernised spelling, used only for searching (see [Search](https://scholaris.joseluissaorin.com/en/knowledge/search.md)). The text that gets cited is never touched.

## Vector spaces

One .spdf can carry vectors from several models at once, each declared in the `espacios` table. The ones Scholaris uses today:

| Space | Model | Dimensions | When |
| --- | --- | --- | --- |
| gemini-embedding-2@1536 | Gemini Embedding 2, multimodal, truncated | 1536 | By default, in the cloud and at home |
| embeddinggemma-2@768 | EmbeddingGemma 2 (text, image, audio and video), on your machine | 768 (or 512, 256, 128) | In the home version offline |
| qwen3-vl-embedding-2b@2048 | Qwen3-VL Embedding 2B, on InferBox | 2048 | In the home version with InferBox, as an extra space |
| qwen3-embedding-0.6b@1024 | Qwen3 Embedding 0.6B, on Workers AI | 1024 | When there is no Gemini key |

When an .spdf is imported, only the vectors missing for your library's space are computed; text and anchors are not read again.

## Opening it offline with Python

The Python SDK includes an SPDF reader that uses only the standard library (gzip and sqlite3) and opens the file read-only:

```sh
pip install scholaris-sdk                 # only needs requests
pip install "scholaris-sdk[vectores]"     # with numpy, for vector search
```

```py
from scholaris.v2 import SPDF

with SPDF.abrir("vigilar.spdf") as s:
    print(s.documento["metadatos"]["titulo"])
    for f in s.buscar("panóptico", k=5):       # FTS5, accent-insensitive
        print(f.cita, f.texto[:80])            # «(Foucault, 1975, p. 23) …»
    print(s.pagina("145").texto)               # by printed page number
```

The reader also exposes `documentos`, `unidades`, `fragmentos`, `secciones`, `espacios`, `blob`, `original` and `vectores(espacio)`, and does similarity search with `buscar_vector(vector, espacio, k)`. The package is `scholaris-sdk` on PyPI and is imported as `scholaris`; it is licensed under the EUPL-1.2.

One caveat: the Python reader searches the FTS5 index as is, without the modernised-spelling layer that Scholaris applies to the query.

## Opening it without Python

```sh
gzip -dc libro.spdf > libro.sqlite
sqlite3 libro.sqlite "SELECT valor FROM spdf WHERE clave = 'spdf_version'"
sqlite3 libro.sqlite "SELECT texto FROM fragmentos_fts WHERE fragmentos_fts MATCH 'panoptico' LIMIT 3"
```

## Export and import

- From the app, every document exports as an .spdf, with or without the original and the vectors. The server embeds up to 24 MB of files; for more, the export runs in the browser.
- A whole library exports as a **.scholaris** package: a ZIP with one .spdf per document, a `manifest.json` (format `scholaris-biblioteca`, version 1) and a `LEEME.txt` with the library's rights. See [Libraries](https://scholaris.joseluissaorin.com/en/knowledge/libraries.md).
- Importing an .spdf (up to 512 MB) reads nothing again.

## What is missing

There is no public validator for version 4 yet. The full specification lives in the code (`packages/spdf/esquema/v4.1.sql`), which will be published with the rest of the source (see [Self-hosting](https://scholaris.joseluissaorin.com/en/knowledge/self-hosting.md)).
