The SPDF 4.1 format, a library you can take with you
An .spdf is a gzip-compressed SQLite database holding the document already read: text, anchors, sections, figures, vectors and provenance. It opens without Scholaris and offline.
Reviewed on This page as Markdown
What an .spdf is
An .spdf file (Scholaris PDF) stores a document already read: not just the original but everything Scholaris got out of it, so it can be searched and cited elsewhere without reading it again or paying for it twice.
Inside it is a gzip-compressed SQLite database (not a ZIP). Any language with SQLite opens it; uncompressed files are accepted too. The current version is 4.1; the version is stored in the spdf table under the key spdf_version. Old SPDFs (versions 1 to 3, from the first Scholaris) are migrated on the fly when opened.
Each exported .spdf holds one document. The same schema is what each user's library uses on the server.
The tables
| Table | What it stores |
|---|---|
| spdf | Key and value: version, creation date, generator, hash of the original |
| documentos | Type, record (JSON), SHA-256 hash, MIME type, size, number of units, duration, title, authors, year and language |
| unidades | The citable units (pages, audio stretches, slides): anchor, text, notes, header and footer, image, confidence, printed page, times and per-word timings |
| secciones | The section tree |
| fragmentos | The passages that are searched and cited: text, context, section, start and end anchors, and the modernised-spelling layer (texto_busqueda) |
| fragmentos_fts | FTS5 full-text index over text, context, section and modernised spelling, accent-insensitive |
| figuras | Figures, plates and video frames, with region, caption and description |
| espacios | Vector spaces: provider, model, version, dimensions, normalisation and modalities |
| vectores | One vector per target (fragment, unit or figure) and space, as little-endian float32 |
| blobs | The original and the images (only in exported files) |
| procedencia | The log of how it was read: each phase, which provider did it, how long it took and when |
Version 4.1 added the texto_busqueda column: a shadow of the text in modernised spelling, used only for searching (see Search). The text that gets cited is never touched.
Vector spaces
One .spdf can carry vectors from several models at once, each declared in the espacios table. The ones Scholaris uses today:
| Space | Model | Dimensions | When |
|---|---|---|---|
| gemini-embedding-2@1536 | Gemini Embedding 2, multimodal, truncated | 1536 | By default, in the cloud and at home |
| qwen3-vl-embedding-2b@2048 | Qwen3-VL Embedding 2B, on InferBox | 2048 | In the home version with a GPU |
| qwen3-embedding-0.6b@1024 | Qwen3 Embedding 0.6B, on Workers AI | 1024 | When there is no Gemini key |
When an .spdf is imported, only the vectors missing for your library's space are computed; text and anchors are not read again.
Opening it offline with Python
The Python SDK includes an SPDF reader that uses only the standard library (gzip and sqlite3) and opens the file read-only:
pip install scholaris-sdk # only needs requests
pip install "scholaris-sdk[vectores]" # with numpy, for vector searchfrom scholaris.v2 import SPDF
with SPDF.abrir("vigilar.spdf") as s:
print(s.documento["metadatos"]["titulo"])
for f in s.buscar("panóptico", k=5): # FTS5, accent-insensitive
print(f.cita, f.texto[:80]) # «(Foucault, 1975, p. 23) …»
print(s.pagina("145").texto) # by printed page numberThe reader also exposes documentos, unidades, fragmentos, secciones, espacios, blob, original and vectores(espacio), and does similarity search with buscar_vector(vector, espacio, k). The package is scholaris-sdk on PyPI and is imported as scholaris; it is licensed under the EUPL-1.2.
One caveat: the Python reader searches the FTS5 index as is, without the modernised-spelling layer that Scholaris applies to the query.
Opening it without Python
gzip -dc libro.spdf > libro.sqlite
sqlite3 libro.sqlite "SELECT valor FROM spdf WHERE clave = 'spdf_version'"
sqlite3 libro.sqlite "SELECT texto FROM fragmentos_fts WHERE fragmentos_fts MATCH 'panoptico' LIMIT 3"Export and import
- From the app, every document exports as an .spdf, with or without the original and the vectors. The server embeds up to 24 MB of files; for more, the export runs in the browser.
- A whole library exports as a .scholaris package: a ZIP with one .spdf per document, a
manifest.json(formatscholaris-biblioteca, version 1) and aLEEME.txtwith the library's rights. See Libraries. - Importing an .spdf (up to 512 MB) reads nothing again.
What is missing
There is no public validator for version 4 yet. The full specification lives in the code (packages/spdf/esquema/v4.1.sql), which will be published with the rest of the source (see Self-hosting).