Skip to the text
Scholaris

The SPDF 4.1 format, a library you can take with you

An .spdf is a gzip-compressed SQLite database holding the document already read: text, anchors, sections, figures, vectors and provenance. It opens without Scholaris and offline.

Reviewed on This page as Markdown

What an .spdf is

An .spdf file (Scholaris PDF) stores a document already read: not just the original but everything Scholaris got out of it, so it can be searched and cited elsewhere without reading it again or paying for it twice.

Inside it is a gzip-compressed SQLite database (not a ZIP). Any language with SQLite opens it; uncompressed files are accepted too. The current version is 4.1; the version is stored in the spdf table under the key spdf_version. Old SPDFs (versions 1 to 3, from the first Scholaris) are migrated on the fly when opened.

Each exported .spdf holds one document. The same schema is what each user's library uses on the server.

The tables

TableWhat it stores
spdfKey and value: version, creation date, generator, hash of the original
documentosType, record (JSON), SHA-256 hash, MIME type, size, number of units, duration, title, authors, year and language
unidadesThe citable units (pages, audio stretches, slides): anchor, text, notes, header and footer, image, confidence, printed page, times and per-word timings
seccionesThe section tree
fragmentosThe passages that are searched and cited: text, context, section, start and end anchors, and the modernised-spelling layer (texto_busqueda)
fragmentos_ftsFTS5 full-text index over text, context, section and modernised spelling, accent-insensitive
figurasFigures, plates and video frames, with region, caption and description
espaciosVector spaces: provider, model, version, dimensions, normalisation and modalities
vectoresOne vector per target (fragment, unit or figure) and space, as little-endian float32
blobsThe original and the images (only in exported files)
procedenciaThe log of how it was read: each phase, which provider did it, how long it took and when

Version 4.1 added the texto_busqueda column: a shadow of the text in modernised spelling, used only for searching (see Search). The text that gets cited is never touched.

Vector spaces

One .spdf can carry vectors from several models at once, each declared in the espacios table. The ones Scholaris uses today:

SpaceModelDimensionsWhen
gemini-embedding-2@1536Gemini Embedding 2, multimodal, truncated1536By default, in the cloud and at home
qwen3-vl-embedding-2b@2048Qwen3-VL Embedding 2B, on InferBox2048In the home version with a GPU
qwen3-embedding-0.6b@1024Qwen3 Embedding 0.6B, on Workers AI1024When there is no Gemini key

When an .spdf is imported, only the vectors missing for your library's space are computed; text and anchors are not read again.

Opening it offline with Python

The Python SDK includes an SPDF reader that uses only the standard library (gzip and sqlite3) and opens the file read-only:

pip install scholaris-sdk                 # only needs requests
pip install "scholaris-sdk[vectores]"     # with numpy, for vector search
from scholaris.v2 import SPDF

with SPDF.abrir("vigilar.spdf") as s:
    print(s.documento["metadatos"]["titulo"])
    for f in s.buscar("panóptico", k=5):       # FTS5, accent-insensitive
        print(f.cita, f.texto[:80])            # «(Foucault, 1975, p. 23) …»
    print(s.pagina("145").texto)               # by printed page number

The reader also exposes documentos, unidades, fragmentos, secciones, espacios, blob, original and vectores(espacio), and does similarity search with buscar_vector(vector, espacio, k). The package is scholaris-sdk on PyPI and is imported as scholaris; it is licensed under the EUPL-1.2.

One caveat: the Python reader searches the FTS5 index as is, without the modernised-spelling layer that Scholaris applies to the query.

Opening it without Python

gzip -dc libro.spdf > libro.sqlite
sqlite3 libro.sqlite "SELECT valor FROM spdf WHERE clave = 'spdf_version'"
sqlite3 libro.sqlite "SELECT texto FROM fragmentos_fts WHERE fragmentos_fts MATCH 'panoptico' LIMIT 3"

Export and import

  • From the app, every document exports as an .spdf, with or without the original and the vectors. The server embeds up to 24 MB of files; for more, the export runs in the browser.
  • A whole library exports as a .scholaris package: a ZIP with one .spdf per document, a manifest.json (format scholaris-biblioteca, version 1) and a LEEME.txt with the library's rights. See Libraries.
  • Importing an .spdf (up to 512 MB) reads nothing again.

What is missing

There is no public validator for version 4 yet. The full specification lives in the code (packages/spdf/esquema/v4.1.sql), which will be published with the rest of the source (see Self-hosting).