# Scholaris: todas las páginas públicas / all public pages Fuente / source: https://scholaris.joseluissaorin.com/llms.txt · Generado / generated: 2026-10-06 # ===== CASTELLANO ===== --- # Todo lo que hay que saber de Scholaris URL: https://scholaris.joseluissaorin.com/saber > Qué lee Scholaris, cómo encuentra y cita, qué formato guarda, qué hace con tus datos, cuánto cuesta y cómo lo usa un agente. Más detallado que la portada y escrito para leerlo despacio. ## Para qué existe esta parte La [portada](https://scholaris.joseluissaorin.com/acerca.md) cuenta por qué existe Scholaris. Estas hojas cuentan cómo funciona, con los datos concretos: los formatos que entran, de dónde sale cada número de página, qué proveedor lee qué, cuánto tarda y cuánto cuesta lo que hemos medido, y lo que todavía no hace. Si algo de lo que pone aquí no coincide con lo que ves en la aplicación, manda la aplicación y esta página está equivocada: escríbenos y la corregimos. Cada hoja tiene su gemelo en Markdown (la misma dirección terminada en `.md`), y hay un índice para máquinas en [/llms.txt](https://scholaris.joseluissaorin.com/llms.txt) y el texto entero de todas las hojas en [/llms-full.txt](https://scholaris.joseluissaorin.com/llms-full.txt). Si eres un agente, empieza por la [hoja para agentes](https://scholaris.joseluissaorin.com/agentes.md). ## Las hojas {#hojas} - [Qué es Scholaris y qué no hará nunca](https://scholaris.joseluissaorin.com/saber/que-es.md): Una biblioteca que lee tus fuentes y, cuando preguntas, señala la página impresa o el segundo exactos. Las reglas que la sostienen: no inventar citas, enseñar la procedencia y dejar el pensamiento a quien escribe. - [Qué formatos entran y cómo se ancla cada cita](https://scholaris.joseluissaorin.com/saber/formatos.md): PDF digitales y escaneados, fotos, EPUB, Word, presentaciones, hojas de cálculo, audio, vídeo, webs, YouTube y pódcast. Para cada uno, cómo se lee y qué ancla queda: folio impreso, segundo, diapositiva, filas o párrafo. - [El formato SPDF 4.1, una biblioteca que te llevas](https://scholaris.joseluissaorin.com/saber/spdf.md): Un .spdf es una base de datos SQLite comprimida con gzip que guarda el documento ya leído: texto, anclas, secciones, figuras, vectores y procedencia. Se abre sin Scholaris y sin conexión. - [Cómo busca, por palabras, por sentido y por imagen](https://scholaris.joseluissaorin.com/saber/busqueda.md): Tres búsquedas a la vez (léxica, semántica y visual) que se funden y se reordenan; frases exactas entre comillas; una capa de grafía modernizada para el castellano antiguo y el latín, y resultados en dos tiempos. - [Citas que se comprueban antes de llegar a ti](https://scholaris.joseluissaorin.com/saber/citas.md): Nueve estilos CSL, BibTeX, RIS y CSL-JSON; una autocita que repasa tu borrador y propone cada referencia con su página; verificación de afirmaciones con lógica temporal, e inserción en .docx sin romper el formato. - [Las herramientas para pensar con una biblioteca entera](https://scholaris.joseluissaorin.com/saber/investigacion.md): El grafo de personas, obras, lugares y conceptos enlazado con Wikidata; el mapa de conceptos; el grafo de citas entre tus libros; los vigilantes que avisan cuando algo cambia y los cuadernos con citas que se vuelven a comprobar. - [Bibliotecas que se llenan de golpe y se comparten](https://scholaris.joseluissaorin.com/saber/bibliotecas.md): Llenar una biblioteca con carpetas, ZIP, listas de enlaces o un BibTeX de Zotero; exportarla como paquete .scholaris; compartirla, seguirla, copiarla o publicarla con un enlace, respetando sus derechos. - [Entrevistas, clases y vídeos citables al segundo](https://scholaris.joseluissaorin.com/saber/reproductor.md): Transcripción palabra a palabra con quién habla, tramos citables de 30 a 60 segundos, fotogramas buscables y un reproductor con la transcripción sincronizada donde se cita seleccionando el texto. - [La API v1 y el servidor MCP](https://scholaris.joseluissaorin.com/saber/api-y-mcp.md): Nueve verbos sobre HTTP con una clave, respuestas en JSON o Markdown, y un servidor MCP con OAuth para Claude, Cursor y otros agentes. Límites, alcances, errores y ejemplos para copiar. - [Scholaris en tu ordenador o en tu propia cuenta de Cloudflare](https://scholaris.joseluissaorin.com/saber/version-local.md): La misma aplicación con SQLite y tu disco: en Node, en Docker, como ejecutable de escritorio o desplegada en tu cuenta de Cloudflare. Qué se queda en tu máquina y qué sigue necesitando la nube. - [Qué pasa con tus datos, proveedor a proveedor](https://scholaris.joseluissaorin.com/saber/privacidad.md): Dónde se guardan tus documentos, qué proveedor de IA lee qué, qué se manda a las bases bibliográficas abiertas, cómo se cifran tus claves y qué se borra cuando borras. - [Planes, cupones y límites](https://scholaris.joseluissaorin.com/saber/planes.md): Un plan gratuito para empezar, Pro para trabajar de verdad y la versión local sin cuotas. Cuánto cabe en cada uno, cómo se cuentan las páginas y los minutos, y cómo funcionan los cupones. - [Lo que hemos medido, con sus condiciones](https://scholaris.joseluissaorin.com/saber/rendimiento.md): Tiempos y costes de lectura, exactitud del folio, calidad de búsqueda, citas inventadas y latencias, tal como salieron en el banco de pruebas del 6 de octubre de 2026, con lo que esas cifras no dicen. - [Preguntas frecuentes](https://scholaris.joseluissaorin.com/saber/preguntas.md): Si inventa citas, si sirve con libros escaneados del siglo XVII o con entrevistas en vídeo, qué pasa con tus datos, si funciona sin conexión y otras dudas, contestadas sin rodeos. - [Glosario de Scholaris](https://scholaris.joseluissaorin.com/saber/glosario.md): Las palabras que usa Scholaris, de folio y ancla a pliego, imprenta, SPDF o vigilante, con lo que significan aquí y, cuando viene a cuento, en la tradición del libro. - [Scholaris al lado de Zotero, Elicit, NotebookLM y los demás](https://scholaris.joseluissaorin.com/saber/alternativas.md): Qué hace cada herramienta, cuándo conviene otra antes que Scholaris y qué hace Scholaris que las demás no se proponen. Sin cifras ajenas que no podamos comprobar. - [Registro de cambios](https://scholaris.joseluissaorin.com/saber/cambios.md): Lo que ha cambiado en Scholaris, de la primera versión a la segunda, contado por hitos y en el orden en que se hicieron. - [Para agentes: cómo leer Scholaris y cómo usarlo en nombre de alguien](https://scholaris.joseluissaorin.com/agentes.md): Qué puede leer un agente en esta web y en qué formato, cómo actuar sobre la biblioteca de una persona con la API v1 o el MCP, con qué permisos y límites, y las reglas para citar sin inventar. ## Scholaris en tres frases - Lee lo que le das (PDF digitales y escaneados, fotos de páginas, EPUB, Word, presentaciones, hojas de cálculo, audio, vídeo, webs y YouTube) y guarda de cada pasaje un ancla: la página impresa, el segundo, la diapositiva o el párrafo. - Cuando buscas o preguntas, te devuelve pasajes con su cita lista para pegar, y cada cita se escribe desde el ancla guardada, nunca desde lo que diga un modelo. - No escribe por ti ni decide qué autor tiene razón: señala el sitio y se aparta. ## Quién lo hace Scholaris lo hace [José Luis Saorín Ferrer](https://joseluissaorin.com), filólogo y programador, en Santa Cruz de Tenerife. El correo es [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- # Qué es Scholaris y qué no hará nunca URL: https://scholaris.joseluissaorin.com/saber/que-es > Una biblioteca que lee tus fuentes y, cuando preguntas, señala la página impresa o el segundo exactos. Las reglas que la sostienen: no inventar citas, enseñar la procedencia y dejar el pensamiento a quien escribe. ## Lo que hace Scholaris es una aplicación web (y una versión para tu ordenador) para quien lee para escribir. Subes tus fuentes: libros escaneados, artículos, tesis, apuntes, presentaciones, entrevistas grabadas, clases, vídeos de YouTube, páginas web. Scholaris las lee una vez, con cuidado, y a partir de ahí puedes: - **buscar** en toda tu biblioteca a la vez, por palabras exactas o por sentido, en cualquier lengua, también en castellano antiguo y en latín; - **preguntar**, y recibir una respuesta corta con notas al pie, cada una con su pasaje; - **citar** en el estilo que pida tu revista, con la página impresa de la edición que tienes delante o el minuto exacto de la grabación; - **verificar** si tu biblioteca respalda una frase de tu borrador; - **explorar** lo que tu biblioteca tiene dentro: personas, obras, lugares y conceptos, y cómo se citan unos libros a otros. ## Las cuatro reglas ### Señalar la página exacta Una cita que no dice dónde está no se puede comprobar. Por eso Scholaris guarda de cada pasaje un **ancla**: el folio impreso que se ve en el papel (no el número que pone el visor de PDF), el segundo de una grabación, la diapositiva, la hoja y las filas de una tabla, o la sección y el párrafo de una web con su fecha de consulta. Cómo se calcula cada ancla está en [Formatos y anclas](https://scholaris.joseluissaorin.com/saber/formatos.md). ### No inventar citas Ninguna cita sale de un modelo de lenguaje. El modelo puede redactar una respuesta, pero solo puede citar los pasajes que se le dan, y cada nota se comprueba contra el texto antes de enseñártela; la referencia (autor, año, página) se escribe desde el ancla guardada. Si no hay pasaje, la respuesta lo dice. En el banco de pruebas de citas, sobre 41 afirmaciones, las citas inventadas fueron cero (ver [Rendimiento](https://scholaris.joseluissaorin.com/saber/rendimiento.md)). ### Enseñar la procedencia Cada dato lleva su origen a la vista: de qué fuente salió la ficha (el propio PDF, Crossref, OpenAlex, Open Library, Wikidata), con qué confianza se leyó un número de página, si se dedujo o se leyó impreso, y quién lo corrigió. Lo que no se sabe se deja vacío en lugar de rellenarlo. ### Dejar el pensamiento a quien escribe Scholaris no escribe ensayos, no resume libros para que no tengas que leerlos y no decide qué cita es la buena. Te lleva al sitio; lo que hagas con lo que encuentres es tuyo. ## Lo que no hace (todavía o nunca) - No busca en internet por ti: trabaja sobre lo que tú le das. Para descubrir literatura nueva hay herramientas mejores (ver [Alternativas](https://scholaris.joseluissaorin.com/saber/alternativas.md)). - No escribe tu texto. La función «citar» devuelve tu propio texto con las citas insertadas, no un texto nuevo. - No garantiza que un número de página impreso sea correcto cuando el libro no lo imprime: en ese caso cita la posición física entre corchetes, «p. [12]», y lo dice. - No funciona del todo sin conexión: incluso la versión local necesita un proveedor en la nube para leer páginas (ver [Versión local](https://scholaris.joseluissaorin.com/saber/version-local.md)). ## Para quién es Para la doctoranda con trescientos PDF y una directora que pide la página; para el profesor que prepara clase con veinte años de subrayados; para la periodista que vuelve a una entrevista de hace dos años buscando una frase; para quien traduce, archiva o estudia, y para cualquiera que haya dicho alguna vez «lo leí en algún sitio». --- # Qué formatos entran y cómo se ancla cada cita URL: https://scholaris.joseluissaorin.com/saber/formatos > PDF digitales y escaneados, fotos, EPUB, Word, presentaciones, hojas de cálculo, audio, vídeo, webs, YouTube y pódcast. Para cada uno, cómo se lee y qué ancla queda: folio impreso, segundo, diapositiva, filas o párrafo. ## Cómo entra un fichero Todo pasa primero por la **imprenta**, el conversor de Scholaris. Reconoce el tipo por los primeros bytes del fichero (después por el tipo MIME y, por último, por la extensión), y lo convierte en tu propio navegador siempre que puede. Lo que el navegador no sabe hacer (decodificar ciertas imágenes, sacar fotogramas de un vídeo, convertir una presentación de Keynote) lo hace un contenedor en el servidor con ffmpeg y LibreOffice. Después, lo que hay que leer con los ojos (páginas escaneadas, fotos, diapositivas) lo lee un modelo de visión, cuatro páginas por petición, y lo que hay que oír lo transcribe un modelo de voz. Qué proveedor hace cada cosa está en [Privacidad](https://scholaris.joseluissaorin.com/saber/privacidad.md). ## Formatos y anclas | Lo que subes | Extensiones | Cómo se lee | Ancla y cómo se cita | | --- | --- | --- | --- | | PDF digital | pdf | La capa de texto, con sus líneas y bloques; cabeceras y pies aparte; se usan las etiquetas de página del propio PDF, su índice y sus metadatos | Página impresa: «p. 23», «pp. 23-24», «p. xiv» | | PDF escaneado o con un OCR antiguo | pdf | Visión: cada página como imagen. Si el PDF trae un OCR viejo de mala calidad, se vuelve a leer | Página impresa, o la física entre corchetes: «p. [12]» | | Fotos de un libro | jpg, png, webp, gif, tiff, heic, bmp, avif | Se ordenan por nombre en orden numérico (IMG_2 antes que IMG_10), se enderezan y se recortan los bordes vacíos | Página impresa de cada foto | | Textos | docx, odt, rtf, html, md, txt | Bloques (títulos, párrafos, listas, citas, tablas, notas) con su ruta de títulos | Sección y párrafo: «Introducción, párr. 4» | | EPUB | epub | El orden de lectura, el índice y, sobre todo, la lista de páginas impresas si el EPUB la trae | «p. 23» si hay lista de páginas; si no, «párr. N» | | Presentaciones | pptx, odp, key | Texto y notas del orador de cada diapositiva; ODP y Keynote se convierten enteras con LibreOffice | «diap. 7» | | Hojas de cálculo | xlsx, xls, ods, csv, tsv | Tablas en trozos de 50 filas, con la cabecera repetida en cada trozo | «Hoja1, filas 2-51» | | Audio | mp3, wav, m4a, ogg, opus, flac, aac | Transcripción palabra a palabra, con quién habla; tramos citables de 30 a 60 segundos que acaban en final de frase | «12:04», «1:02:03», «12:04-12:40» | | Vídeo | mp4, mov, webm, mkv, avi | El audio como arriba; además, fotogramas en cada cambio de escena (o cada 20 segundos) que también se pueden buscar | Igual que el audio | | Una web | URL | Se guarda una copia fechada y se parte el artículo en bloques | Sección y párrafo, con la fecha de consulta | | YouTube | URL | No se descarga: Gemini ve el vídeo desde la dirección, en tramos de diez minutos | El segundo | | Vimeo, pódcast, enlaces a audio | URL | Vimeo, el fichero progresivo más pequeño; un pódcast, el audio de su RSS | El segundo | | SPDF y paquetes .scholaris | spdf, scholaris | No se vuelve a leer nada: se copian el texto, las anclas y los vectores | Las anclas que traía | ## El folio impreso El número que importa en una nota al pie es el que está impreso en el papel, no la posición de la página en el fichero. Un libro puede empezar en romanos, saltarse láminas sin numerar o tener dos páginas por imagen. Para cada página, Scholaris junta varias pruebas: 1. las etiquetas de página del propio PDF, si las tiene; 2. los números candidatos de la cabecera, del pie y de los bordes del texto, y lo que vio el lector de visión; 3. la secuencia coherente más larga de esas lecturas; 4. en los casos dudosos, un juez (Jev, de TypeSafe) que decide entre candidatos; 5. y, para las páginas sin número visible, interpolación dentro de cada tramo. Reconoce números romanos (con la página en que se pasa a arábigos), foliación por hojas («23r», «23v»), escaneos a doble página, láminas sin numerar y años que no son folios. Cada página guarda de dónde salió su número (leído, deducido, del EPUB o ninguno) y con qué confianza. **Si un libro no imprime ningún número, no se inventa ninguno:** se cita la posición física entre corchetes. En el banco de pruebas, los 64 folios comprobados a ojo salieron exactos (ver [Rendimiento](https://scholaris.joseluissaorin.com/saber/rendimiento.md)). ## La ortografía original se respeta El lector transcribe con fidelidad: no moderniza «dixo», «assi» o «muger», mantiene u/v, i/j/y y ç, no corrige erratas y deja fuera del cuerpo los titulillos, los reclamos y las signaturas. La única excepción es la s larga (ſ), que se escribe «s». Para que esas grafías se encuentren al buscar existe una capa aparte, explicada en [Búsqueda](https://scholaris.joseluissaorin.com/saber/busqueda.md). ## Límites de tamaño | Plan | Fichero más grande | Almacenamiento | | --- | --- | --- | | Gratis | 200 MB | 1 GB | | Pro | 4 GB | 100 GB | | Versión local | 16 GB | el de tu disco | Por la API v1 en la nube caben ficheros de hasta 95 MB en una sola petición; para más, la aplicación y el SDK de Python suben por partes. No hay una duración máxima para audio y vídeo más allá del tamaño del fichero. ## Lo que todavía no va bien - Los .doc antiguos de Word no tienen un lector propio: conviértelos a .docx antes de subirlos. - Vimeo solo funciona cuando el vídeo ofrece un fichero descargable; si no, bájalo y súbelo. - Las anclas de una web son de párrafo: si la página cambia después, la copia fechada que guardamos es la que manda. --- # El formato SPDF 4.1, una biblioteca que te llevas URL: https://scholaris.joseluissaorin.com/saber/spdf > Un .spdf es una base de datos SQLite comprimida con gzip que guarda el documento ya leído: texto, anclas, secciones, figuras, vectores y procedencia. Se abre sin Scholaris y sin conexión. ## Qué es un .spdf Un fichero **.spdf** (Scholaris PDF) guarda un documento ya leído: no solo el original, sino todo lo que Scholaris sacó de él, de manera que se pueda buscar y citar en otro sitio sin volver a leerlo ni pagar otra vez por ello. Por dentro es una base de datos **SQLite comprimida con gzip** (no es un ZIP). Cualquier lenguaje con SQLite la abre; también se acepta sin comprimir. La versión actual es la **4.1**; la versión se guarda en la tabla `spdf`, con la clave `spdf_version`. Los SPDF antiguos (versiones 1 a 3, de la primera Scholaris) se migran solos al abrirlos. Cada .spdf exportado lleva un documento. El mismo esquema es el que usa, en el servidor, la biblioteca de cada usuario. ## Las tablas | Tabla | Qué guarda | | --- | --- | | spdf | Clave y valor: versión, fecha de creación, programa que lo generó, huella del original | | documentos | Tipo, ficha (JSON), huella SHA-256, tipo MIME, tamaño, número de unidades, duración, título, autores, año e idioma | | unidades | Las unidades citables (páginas, tramos de audio, diapositivas): ancla, texto, notas, cabecera y pie, imagen, confianza, página impresa, tiempos y los tiempos de cada palabra | | secciones | El árbol de secciones | | fragmentos | Los pasajes que se buscan y se citan: texto, contexto, sección, ancla de inicio y de fin, y la capa de grafía modernizada (texto_busqueda) | | fragmentos_fts | Índice de texto completo FTS5 sobre texto, contexto, sección y grafía modernizada, sin distinguir tildes | | figuras | Figuras, láminas y fotogramas, con su región, pie y descripción | | espacios | Los espacios vectoriales: proveedor, modelo, versión, dimensiones, normalización y modalidades | | vectores | Un vector por objetivo (fragmento, unidad o figura) y espacio, en float32 little-endian | | blobs | El original y las imágenes (solo en los ficheros exportados) | | procedencia | El registro de cómo se leyó: cada fase, qué proveedor la hizo, cuánto tardó y cuándo | La versión 4.1 añadió la columna `texto_busqueda`: una sombra del texto con la ortografía modernizada, solo para buscar (ver [Búsqueda](https://scholaris.joseluissaorin.com/saber/busqueda.md)). El texto que se cita no se toca nunca. ## Espacios vectoriales Un mismo .spdf puede llevar vectores de varios modelos a la vez, y cada uno se declara en la tabla `espacios`. Los que usa Scholaris hoy: | Espacio | Modelo | Dimensiones | Cuándo | | --- | --- | --- | --- | | gemini-embedding-2@1536 | Gemini Embedding 2, multimodal, recortado | 1536 | Por defecto, en la nube y en casa | | qwen3-vl-embedding-2b@2048 | Qwen3-VL Embedding 2B, en InferBox | 2048 | En la versión local con GPU | | qwen3-embedding-0.6b@1024 | Qwen3 Embedding 0.6B, en Workers AI | 1024 | Si no hay clave de Gemini | Al importar un .spdf solo se calculan los vectores que falten para el espacio que use tu biblioteca; el texto y las anclas no se vuelven a leer. ## Abrirlo sin conexión con Python El SDK de Python trae un lector de SPDF que solo usa la biblioteca estándar (gzip y sqlite3) y abre el fichero en modo de solo lectura: ```sh pip install scholaris-sdk # solo depende de requests pip install "scholaris-sdk[vectores]" # con numpy, para buscar por vectores ``` ```py from scholaris.v2 import SPDF with SPDF.abrir("vigilar.spdf") as s: print(s.documento["metadatos"]["titulo"]) for f in s.buscar("panóptico", k=5): # FTS5, insensible a acentos print(f.cita, f.texto[:80]) # «(Foucault, 1975, p. 23) …» print(s.pagina("145").texto) # por número de página impreso ``` El lector también da acceso a `documentos`, `unidades`, `fragmentos`, `secciones`, `espacios`, `blob`, `original` y `vectores(espacio)`, y busca por similitud con `buscar_vector(vector, espacio, k)`. El paquete se llama `scholaris-sdk` en PyPI y se importa como `scholaris`; tiene licencia EUPL-1.2. Una advertencia: la búsqueda del lector de Python usa el índice FTS5 tal cual, sin la capa de grafía modernizada que aplica Scholaris a la consulta. ## Abrirlo sin Python ```sh gzip -dc libro.spdf > libro.sqlite sqlite3 libro.sqlite "SELECT valor FROM spdf WHERE clave = 'spdf_version'" sqlite3 libro.sqlite "SELECT texto FROM fragmentos_fts WHERE fragmentos_fts MATCH 'panoptico' LIMIT 3" ``` ## Exportar e importar - Desde la aplicación, cada documento se exporta como .spdf, con o sin el original y los vectores. Desde el servidor caben hasta 24 MB de ficheros incrustados; para más, la exportación se hace en el navegador. - Una biblioteca entera se exporta como paquete **.scholaris**: un ZIP con un .spdf por documento, un `manifest.json` (formato `scholaris-biblioteca`, versión 1) y un `LEEME.txt` con los derechos de la biblioteca. Ver [Bibliotecas](https://scholaris.joseluissaorin.com/saber/bibliotecas.md). - Importar un .spdf (hasta 512 MB) no vuelve a leer nada. ## Lo que falta Aún no hay un validador público para la versión 4. La especificación completa vive en el código (`packages/spdf/esquema/v4.1.sql`), que se publicará con el resto del código fuente (ver [Versión local](https://scholaris.joseluissaorin.com/saber/version-local.md)). --- # Cómo busca, por palabras, por sentido y por imagen URL: https://scholaris.joseluissaorin.com/saber/busqueda > Tres búsquedas a la vez (léxica, semántica y visual) que se funden y se reordenan; frases exactas entre comillas; una capa de grafía modernizada para el castellano antiguo y el latín, y resultados en dos tiempos. ## Tres caminos que se juntan Cada consulta va a la vez por tres caminos y los resultados se funden por rango recíproco (RRF), con pesos que dependen de lo que parece que buscas: | Camino | Qué encuentra | Cómo | | --- | --- | --- | | Léxico | Las palabras que escribes, sin distinguir tildes | Índice FTS5 de SQLite con BM25; los títulos de sección y el contexto pesan menos que el texto | | Semántico | Lo que significa la consulta, aunque el pasaje use otras palabras o esté en otra lengua | Vectores de Gemini Embedding 2 (1536 dimensiones) en Cloudflare Vectorize | | Visual | Láminas, figuras, gráficos y fotogramas | Los mismos vectores multimodales, aplicados a las imágenes de las páginas y a las figuras | | Si la consulta parece… | léxico | semántico | visual | | --- | --- | --- | --- | | una idea | 0,35 | 1 | 0 | | una cita literal | 1 | 0,45 | 0 | | algo que se ve | 0,35 | 0,6 | 1 | | una pregunta por fechas | 0,35 | 1 | 0 | Los treinta primeros candidatos los reordena un juez, Jev (de TypeSafe), que contesta para cada pasaje si responde a la consulta; la puntuación final es un 20 % la de la fusión y un 80 % la del juez. El camino visual solo pesa cuando buscas algo que se ve: medido solo, acierta poco con preguntas de texto y las empeoraba. ## Frases exactas Entre comillas («…», "…" o “…”) la frase se busca literal, sin modelo de por medio. Primero salen las coincidencias exactas y después los pasajes cercanos en sentido. Si la frase no aparece en ningún sitio, Scholaris lo dice («La frase exacta no aparece; se busca por sentido») en lugar de fingir que la encontró. ## Castellano antiguo y latín {#grafia} Un libro del siglo XVII escribe «dixo», «muger», «assi» o «agora»; tú buscas «dijo», «mujer», «así», «ahora». Para que se encuentren sin tocar el texto que se cita, cada pasaje lleva una **sombra en grafía modernizada** que solo sirve para buscar, y la consulta se reduce con las mismas reglas: - **castellano de los siglos XVI a XVIII:** listas de palabras (fee → fe, agora → ahora, mesmo → mismo, truxo → trajo) y reglas fonéticas y ortográficas (b/v, c/z, g/j, h, x → j, ph, rr…); - **latín:** u/v, i/j, æ/œ, michi → mihi, enclíticos -que/-ne/-ve y el lematizador de Schinke y otros (1996); - **francés e italiano antiguos:** reglas básicas. La capa se aplica al latín siempre y a las demás lenguas cuando el documento es anterior a una fecha de corte (1830 para el castellano, 1800 para el francés y el italiano) o el propio texto da señales (ſ, ç, «assi», «dixo»…). Son reglas, sin modelos. En una comedia de Lope, la exhaustividad media por consulta pasó del 32 % al 100 %; en documentos modernos, las búsquedas devuelven exactamente lo mismo que sin la capa (ver [Rendimiento](https://scholaris.joseluissaorin.com/saber/rendimiento.md)). ## Resultados en dos tiempos En la aplicación, la búsqueda llega en tres entregas por un canal de eventos: primero lo que encuentra el índice léxico, sin esperar al vector de la consulta; después la fusión de los tres caminos, hacia los 350 ms; y al final el orden del juez, hacia los 650 ms de mediana medidos desde casa. Lo primero que ves ya es útil y no se reordena de golpe. ## Filtros Se puede filtrar por biblioteca, documento, tipo, autor, idioma y año (desde, hasta). Algunos filtros se entienden escritos en la propia consulta: «antes de 1980», «desde 1950», el apellido de un autor de tu biblioteca. «Ir a la página 145» salta a esa página impresa. ## Lo que devuelve cada resultado El pasaje literal, su cita lista para pegar («(Foucault, 1975, p. 23)»), el localizador exacto, el ancla y un enlace que abre el lector en esa página o en ese segundo. La búsqueda nunca escribe una cita: la cita sale del ancla. ## También - **Preguntar:** una respuesta breve en Markdown con notas al pie; el redactor solo puede citar los pasajes que se le dan y cada nota se comprueba. - **Parecidos:** pasajes y figuras parecidos a uno dado. - **En varias lenguas:** la consulta traducida a las lenguas de tu biblioteca. - **Bibliotecas que sigues:** buscar a la vez en las tuyas y en las que otros comparten contigo. --- # Citas que se comprueban antes de llegar a ti URL: https://scholaris.joseluissaorin.com/saber/citas > Nueve estilos CSL, BibTeX, RIS y CSL-JSON; una autocita que repasa tu borrador y propone cada referencia con su página; verificación de afirmaciones con lógica temporal, e inserción en .docx sin romper el formato. ## Subir y citar, paso a paso {#subir-y-citar} 1. Entra en Scholaris, o pruébalo sin cuenta desde [la demostración](https://scholaris.joseluissaorin.com/?demostracion). 2. En la Biblioteca, arrastra el fichero (o pega una dirección: una web, un PDF, YouTube, un pódcast). Las primeras páginas de un PDF digital se pueden buscar en menos de un segundo; las de un escaneado, en unos segundos. 3. En Buscar, escribe la idea con tus palabras o la frase exacta entre comillas. 4. En el resultado, copia la referencia: sale en tu estilo, con la página impresa o el minuto, y como texto con formato (cursivas incluidas). 5. Para un texto entero, ve a Escribir y pega tu borrador o sube tu .docx: la autocita propone una cita para cada afirmación, con su página, y tú aceptas o descartas una a una. 6. Exporta el texto citado (.docx, Markdown, texto o LaTeX) y la bibliografía (BibTeX, RIS o CSL-JSON). ## Estilos Las citas las da formato citeproc-js con los ficheros CSL oficiales, en cuatro locales (es-ES, en-US, fr-FR, it-IT). El estilo por defecto es APA. | Identificador | Estilo | | --- | --- | | apa | APA, 7.ª edición | | chicago-author-date | Chicago, 18.ª edición, autor-fecha | | chicago-note-bibliography | Chicago, 18.ª edición, notas y bibliografía | | mla | MLA, 9.ª edición | | harvard | Harvard (Cite Them Right, 12.ª edición) | | iso690 | ISO 690 autor-fecha, en castellano | | iso690-en | ISO 690 autor-fecha, en inglés | | iso690-numerico | ISO 690 numérico | | ieee | IEEE | Cuando la fecha de la obra original no coincide con la de la edición, sale como fecha original («1605/2004»); la edición se indica salvo que sea una reimpresión del mismo año. ## De dónde sale cada cita La cita no la escribe un modelo. El localizador («p. 23», «12:04», «diap. 7») sale del ancla guardada al leer el documento (ver [Formatos y anclas](https://scholaris.joseluissaorin.com/saber/formatos.md)), y el autor, el año y el título salen de la ficha, cuya procedencia se puede consultar campo a campo. La ficha se contrasta con Crossref, OpenAlex, Open Library y Wikidata, y solo se acepta un registro externo si el título coincide y además coincide el autor o el año. ## La autocita Para un borrador entero, Scholaris: 1. lo parte en párrafos y afirmaciones; 2. recupera catorce pasajes candidatos para cada una; 3. deja que un modelo proponga qué pasaje respalda cada afirmación; 4. comprueba en código que la evidencia aparece literal en el pasaje, que no se cita una afirmación negativa como si fuera positiva, que están los términos clave y que la cronología es posible; 5. y pasa cada propuesta por un juez (Jev, de TypeSafe) en lotes de ocho. Las propuestas con probabilidad de 0,7 o más quedan aceptadas; entre 0,4 y 0,7, para revisar; por debajo, se descartan. Las páginas consecutivas verificadas se juntan («pp. 23-25»). En el banco de pruebas, la precisión fue del 95,1 %, la exhaustividad del 96,7 % y las citas inventadas, cero. ## Verificar una afirmación Le das una frase y Scholaris dice si tu biblioteca la respalda, con los pasajes: respaldada : el mejor pasaje la apoya directamente, con probabilidad de 0,7 o más. parcial : algún pasaje la apoya en parte (probabilidad de 0,4 o más). contradicha : algún pasaje dice lo contrario, con probabilidad de 0,5 o más. sin_respaldo : no hay pasajes que la apoyen. El juez clasifica además la relación de cada pasaje con la afirmación: apoyo directo, aplicación de un marco, contexto, contradicción, opinión referida o afirmación negativa. ## La lógica temporal Comparar fechas no se le pide a un modelo: lo hace el código. - Se usa el año de la obra original, no el de la edición. - Un texto no puede apoyar directamente una afirmación sobre algo que se inventó después: Scholaris conoce el año de aparición de términos como «internet» (1983), «deconstrucción» (1967) o «transformer» (2017). - Si la afirmación dice que un autor «anticipó» o «se basó en» otro, la cronología tiene que ser posible; si no lo es, la cita se descarta. - Si la afirmación habla de un año más de veinte años posterior a la fuente, el apoyo directo baja a «aplicación de un marco». - Una fuente posterior al año de tu propio texto no puede citarse en él. - Sin fecha («s. f.»), no se aplica nada de esto. ## Dentro de Word Scholaris abre el .docx, inserta cada cita en su sitio respetando el formato de cada fragmento de texto, usa notas al pie de verdad de Word en los estilos de notas y añade la bibliografía al final. También acepta ODT, texto, Markdown y HTML, hasta 50 MB. ## Copiar una referencia Cada documento tiene un botón para copiar su referencia en tu estilo (o en otro, con vista previa). Se copia como texto con formato, con las cursivas, y como texto plano. ## Exportar e importar - Exportar: BibTeX, RIS y CSL-JSON; la bibliografía, como texto, HTML o Markdown. - Importar un BibTeX (por ejemplo, el de Zotero): cada entrada se casa con tus documentos por DOI, después por ISBN y después por título; las que no casan quedan como referencias sin texto. --- # Las herramientas para pensar con una biblioteca entera URL: https://scholaris.joseluissaorin.com/saber/investigacion > El grafo de personas, obras, lugares y conceptos enlazado con Wikidata; el mapa de conceptos; el grafo de citas entre tus libros; los vigilantes que avisan cuando algo cambia y los cuadernos con citas que se vuelven a comprobar. ## Personas, obras, lugares y conceptos Scholaris lee toda la biblioteca buscando entidades: personas, obras, lugares, organizaciones, conceptos, acontecimientos y fechas (y distingue a los personajes de ficción). Un modelo las propone por lotes de unos 48 000 caracteres, y después el código busca cada mención en el texto, de modo que cada aparición tiene su ancla exacta y se puede abrir en su página. Cada entidad se enlaza, cuando se puede, con su ficha de Wikidata. Dos entidades quedan unidas cuando aparecen cerca: el peso de la arista baja con la distancia dentro del pasaje y es menor entre pasajes vecinos. Para las relaciones más fuertes de cada documento, un modelo pone nombre a la relación. Desde el grafo se puede ver a los vecinos de una entidad, su línea temporal y el camino más corto entre dos. Medido en el banco, extraer las entidades de 300 páginas cuesta unos 0,03 $. ## El mapa de conceptos Una vista de la biblioteca como un mapa del cielo: los vectores de los pasajes y las figuras se agrupan, se proyectan a dos dimensiones con UMAP (sobre una muestra de hasta 8000 puntos) y cada grupo recibe un nombre. Sirve para ver de qué habla tu biblioteca y qué queda lejos de qué. Aparte está **Conceptos**: escribes un concepto y Scholaris reúne dónde se define, dónde se aplica, dónde se critica y dónde solo se menciona, con su pasaje y su página. Se exporta en CSV, JSON, JSONL, HTML, BibTeX, TEI y XLSX. ## El grafo de citas Cómo se citan unos documentos de tu biblioteca a otros. Las entradas de la bibliografía de cada documento se resuelven contra tu biblioteca por DOI, por título (con el año compatible) o por primer autor y año cuando solo hay un candidato, y cada enlace lleva su confianza. Lo que tus libros citan y tú no tienes aparece como **referencias huérfanas**: una lista de lo que te falta por leer. El grafo solo mira dentro de tu biblioteca; no consulta bases externas. ## Vigilantes Un vigilante es una búsqueda guardada que se repite sola: a mano, al subir un documento, cada día o cada semana (en la nube, cada día a las 06:17 UTC, y los lunes los semanales). Te avisa cuando aparece un documento nuevo entre los primeros resultados o cuando la respuesta cambia: si sube o baja la confianza, o si dice otra cosa. ## Cuadernos Notas en Markdown con fichas: un pasaje, una página, una figura, un tramo de audio, una nota tuya o una síntesis. Cada ficha guarda una copia verificada del pasaje, su ancla y una huella, y con **volver a verificar** Scholaris comprueba que la biblioteca sigue diciendo lo mismo: citas vivas. La síntesis de un cuaderno solo puede citar las fichas del propio cuaderno; cualquier otra marca se elimina. ## Perspectivas Documentos que llevas tiempo sin abrir, huecos (temas de los que tienes poco) y recomendaciones dentro de lo que ya tienes. --- # Bibliotecas que se llenan de golpe y se comparten URL: https://scholaris.joseluissaorin.com/saber/bibliotecas > Llenar una biblioteca con carpetas, ZIP, listas de enlaces o un BibTeX de Zotero; exportarla como paquete .scholaris; compartirla, seguirla, copiarla o publicarla con un enlace, respetando sus derechos. ## Llenar una biblioteca de una vez Desde «Llenar biblioteca» se puede soltar una carpeta entera, muchos ficheros, un ZIP, una lista de enlaces, un BibTeX o RIS con la carpeta de Zotero al lado, o un montón de .spdf. Antes de empezar, Scholaris enseña cuánto va a tardar y cuánto va a costar, y deja elegir entre dos modos: | Modo | Qué hace | Para qué | | --- | --- | --- | | Rápido | Lee en el momento | Cuando lo necesitas hoy | | Económico | Las páginas difíciles van por el lote de Gemini, a mitad de precio, y las fáciles por un lector más barato; tarda horas | Para llenar una biblioteca grande | Los ficheros repetidos (misma huella SHA-256) se detectan y no se leen dos veces. El lote se puede pausar y reanudar. ## Paquetes .scholaris Una biblioteca entera se exporta como un fichero **.scholaris**: un ZIP con un [.spdf](https://scholaris.joseluissaorin.com/saber/spdf.md) por documento, un `manifest.json` con la biblioteca, sus derechos y sus documentos, y un `LEEME.txt`. Se puede incluir o no los originales y los vectores. Importarlo en otra cuenta, o en la versión local, no vuelve a leer nada. ## Compartir - **Invitar** por correo, con permiso de lectura, de edición o de administración, y con caducidad (de un día a diez años). - **Seguir** una biblioteca que te han compartido, y buscar en ella junto a las tuyas. - **Copiar** a tu biblioteca: la copia es por referencia, así que no se duplican los ficheros ni se vuelve a leer nada. - **Enlace público** de solo lectura, con contraseña opcional, caducidad, contador de visitas y posibilidad de revocarlo. Las páginas de esos enlaces piden a los buscadores que no las indexen. ## Derechos Cada biblioteca declara sus derechos: sin indicar, dominio público, CC0, CC BY, CC BY-SA, CC BY-NC, CC BY-NC-SA, uso privado o con permiso. Para publicar con un enlace una biblioteca que no es abierta hay que confirmar que es para uso privado, y el paquete de una biblioteca no abierta lleva un aviso de no redistribuir. Compartir un libro con derechos con quien no debe es responsabilidad de quien lo comparte; Scholaris solo te pregunta antes. --- # Entrevistas, clases y vídeos citables al segundo URL: https://scholaris.joseluissaorin.com/saber/reproductor > Transcripción palabra a palabra con quién habla, tramos citables de 30 a 60 segundos, fotogramas buscables y un reproductor con la transcripción sincronizada donde se cita seleccionando el texto. ## La transcripción El audio se transcribe en tramos de diez minutos con dos segundos de solape. Cuando hay clave de Gemini, transcribe Gemini Transcribe, que distingue hablantes; si falla, Whisper (large-v3-turbo, en Workers AI). Después, un modelo identifica el reparto (quién habla y en qué papel: entrevistador, invitado, presentadora) y asigna cada frase a una persona con nombre, lo que corrige las etiquetas que no coinciden de un tramo a otro. En dos entrevistas del banco, la atribución fue correcta en 20 de 21 turnos y en 23 de 23. El texto se parte en **tramos citables** de 30 a 60 segundos (45 de media) que acaban en final de frase. Cada tramo es una unidad con su ancla de tiempo y su hablante, y se guardan además los tiempos de cada palabra. ## Los fotogramas En los vídeos se toma un fotograma en cada cambio de escena y, si no hay cambios, uno cada 20 segundos (nunca más de uno cada 10). Los fotogramas se describen y se vectorizan, de modo que se puede buscar «la pizarra con el diagrama» y llegar al segundo exacto. ## YouTube, Vimeo y pódcast Un vídeo de YouTube no se descarga: Gemini lo ve desde su dirección, a trozos de diez minutos, y el título y el canal salen de su ficha pública. De Vimeo se toma el fichero descargable más pequeño; de un pódcast, el audio que anuncia su RSS. ## El reproductor - Línea de tiempo con los turnos de cada hablante, los capítulos y una vista previa del fotograma al pasar por encima. - Transcripción sincronizada que resalta la palabra que suena; un clic en una palabra salta a ese momento. - Para citar, se selecciona el texto: la cita sale con su intervalo exacto («12:04-12:40»). - Un reproductor pequeño sigue sonando mientras navegas por el resto de la aplicación. ## Lo que hemos medido Una entrevista en vídeo de 54 minutos quedó lista en 45 segundos (0,41 $), y una de dos horas, en 70 segundos (0,94 $); transcribir cuesta unos 0,005 $ por minuto. En la primera Scholaris, la misma entrevista de 54 minutos tardaba tres horas y 39 minutos. Todas las cifras y sus condiciones están en [Rendimiento](https://scholaris.joseluissaorin.com/saber/rendimiento.md). ## Lo que falta Los tiempos de cada palabra se estiman dentro de cada segmento, no se miden uno a uno. Algunos vídeos con códecs poco comunes aún no se convierten a un formato que todos los navegadores reproduzcan. --- # La API v1 y el servidor MCP URL: https://scholaris.joseluissaorin.com/saber/api-y-mcp > Nueve verbos sobre HTTP con una clave, respuestas en JSON o Markdown, y un servidor MCP con OAuth para Claude, Cursor y otros agentes. Límites, alcances, errores y ejemplos para copiar. ## En una frase Todo lo que hace la aplicación con tus documentos se puede hacer desde fuera con una cabecera `Authorization: Bearer sch_…` y nueve verbos. La guía completa, con una sesión grabada de verdad, está en [/api](https://scholaris.joseluissaorin.com/api.md); la especificación, en [OpenAPI 3.1](https://scholaris.joseluissaorin.com/api/v1/openapi.json); y las instrucciones para modelos, en [/api/v1/llms.txt](https://scholaris.joseluissaorin.com/api/v1/llms.txt). ## Los verbos Base: `https://scholaris.joseluissaorin.com/api/v1` | Método | Ruta | Qué hace | | --- | --- | --- | | POST | /documentos | Sube un fichero (cuerpo crudo, multipart con «archivo» o JSON con «url») | | GET | /documentos | Lista la biblioteca (q, estado, cursor; hasta 200 por página) | | GET | /documentos/{id} | La ficha; con esperar=30 espera a que esté leído | | DELETE | /documentos/{id} | Lo borra con todo lo derivado | | GET | /documentos/{id}/texto | El texto por página impresa, [página física] o tramo de tiempo | | GET, POST | /buscar | Busca pasajes (k hasta 50) | | POST | /preguntar | Responde con notas comprobadas; por eventos con stream | | POST | /citar | Tu texto (o tu .docx) con las citas y la bibliografía | | POST | /verificar | ¿Respalda tu biblioteca una afirmación? | Cualquier lectura admite `formato=markdown` o la cabecera `Accept: text/markdown`. Lo que tarda (subir y citar) espera por defecto; con `esperar=0` o `Prefer: respond-async` responde 202 con `progreso_url`. ```sh export SCHOLARIS=sch_… # Ajustes → Claves de API # Subir (espera a que esté leído, hasta 60 s; si tarda más, 202 y progreso_url) curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos?nombre=articulo.pdf" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/pdf" --data-binary @articulo.pdf # Buscar, en Markdown curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=atención escalada" -d k=5 -d formato=markdown # Verificar una afirmación curl -s https://scholaris.joseluissaorin.com/api/v1/verificar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"afirmacion": "El Transformer prescinde de la recurrencia."}' ``` ```py from scholaris.api import Scholaris s = Scholaris("sch_…") # o la variable SCHOLARIS_CLAVE doc = s.subir("articulo.pdf") # también una URL for p in s.buscar("atención escalada", k=3): print(p["cita"], p["texto"][:80], p["enlace"]) ``` ## Claves y alcances Las claves se crean en Ajustes → Claves de API y se enseñan una sola vez; Scholaris solo guarda su huella SHA-256. Pueden caducar al cabo de los días que elijas. Cada clave tiene alcances: lectura : buscar, leer, preguntar y verificar. escritura : subir y borrar documentos, y citar un texto entero (la autocita guarda su trabajo). mcp : usar la clave en el servidor MCP. Sin alcances explícitos, una clave nueva tiene lectura y mcp. ## Límites | | Gratis | Pro | | --- | --- | --- | | Peticiones por minuto | 120 | 600 | | Búsquedas al día | 100 | 5000 | | Autocitas al mes | 5 | 500 | | Páginas o minutos leídos al mes | 1500 | 60 000 | | Fichero por petición en la API v1 | 95 MB | 95 MB | Al pasarse del ritmo se recibe un 429 con `Retry-After`; al agotar una cuota, un 402 (`cuota_superada` o `requiere_pro`). Para reintentar sin duplicar, `Idempotency-Key` en subir y citar: con la misma clave, durante 24 horas, se devuelve el mismo recurso. Un fichero idéntico (misma huella) devuelve el que ya había, con `duplicado: true`. ## Errores Todos tienen la misma forma, con el mensaje en castellano y en inglés: ```json { "error": { "codigo": "prohibido", "mensaje": "Esta clave de API es de solo lectura.", "message": "This key is not allowed to do that.", "estado": 403, "documentacion": "https://scholaris.joseluissaorin.com/api#errores" } } ``` Códigos: `no_autenticado` (401), `prohibido` (403), `peticion_invalida` (400), `no_encontrado` (404), `conflicto` (409), `demasiado_grande` (413), `cuota_superada` y `requiere_pro` (402), `limite_de_ritmo` (429), `proveedor_fallo` (502), `no_disponible` e `interno`. ## El servidor MCP En `https://scholaris.joseluissaorin.com/mcp`, por HTTP («Streamable HTTP», sin estado). Ofrece cuatro herramientas: | Herramienta | Qué hace | | --- | --- | | search | Búsqueda híbrida; pasajes con su localizador exacto | | cite | Cita CSL de un fragmento o de un documento, en el estilo y la lengua que se pidan | | open_page | El texto entero de una página, por posición física o por folio impreso | | verify_claim | Veredicto sobre una afirmación, con los pasajes que la apoyan o la contradicen | Para conectarse hay dos caminos: - **OAuth 2.1**, para Claude (web y escritorio) y cualquier cliente que lo hable: Ajustes → Conectores → Añadir conector personalizado, con la dirección `https://scholaris.joseluissaorin.com/mcp`. El cliente se registra solo, tú entras en Scholaris y concedes acceso. Ese acceso es de solo lectura: buscar, abrir páginas, citar y verificar; no puede subir, cambiar ni borrar. Los tokens duran una hora y se renuevan. - **Una clave** con el alcance `mcp`, para Claude Code, Cursor, Windsurf y los demás: ```sh claude mcp add --transport http scholaris https://scholaris.joseluissaorin.com/mcp \ --header "Authorization: Bearer sch_…" ``` ```json { "mcpServers": { "scholaris": { "url": "https://scholaris.joseluissaorin.com/mcp", "headers": { "Authorization": "Bearer sch_…" } } } } ``` Los metadatos de OAuth están donde los buscan los clientes: [/.well-known/oauth-authorization-server](https://scholaris.joseluissaorin.com/.well-known/oauth-authorization-server) y [/.well-known/oauth-protected-resource/mcp](https://scholaris.joseluissaorin.com/.well-known/oauth-protected-resource/mcp). Hay además una tarjeta del servidor en [/.well-known/mcp/server-card.json](https://scholaris.joseluissaorin.com/.well-known/mcp/server-card.json). ## Para agentes Las reglas de uso (citar solo lo que devuelve la API, copiar la cita tal cual, dar el enlace, decir cuándo no se encuentra nada) están en la [hoja para agentes](https://scholaris.joseluissaorin.com/agentes.md). --- # Scholaris en tu ordenador o en tu propia cuenta de Cloudflare URL: https://scholaris.joseluissaorin.com/saber/version-local > La misma aplicación con SQLite y tu disco: en Node, en Docker, como ejecutable de escritorio o desplegada en tu cuenta de Cloudflare. Qué se queda en tu máquina y qué sigue necesitando la nube. ## Cuatro maneras La instancia de [scholaris.joseluissaorin.com](https://scholaris.joseluissaorin.com/acerca.md) es la versión alojada, la que se paga con el plan Pro. El mismo código funciona también en tu ordenador: | Cómo | Qué es | Dónde quedan tus datos | | --- | --- | --- | | Node | La misma API sobre Node, SQLite (con sqlite-vec para los vectores) y el disco, con una cola que se reanuda si se corta; escucha en el puerto 8790 | La carpeta que elijas | | Docker | Lo mismo en un contenedor con ffmpeg; una variante añade InferBox, un servidor de modelos con GPU NVIDIA | Un volumen de Docker | | Escritorio | Un solo ejecutable (Bun) para macOS, Windows y Linux, con la web dentro, que abre el navegador al arrancar | ~/Scholaris | | Tu cuenta de Cloudflare | Un guion crea la base de datos, el almacenamiento, el índice de vectores y la cola, y despliega el Worker (necesita Workers de pago) | Tu cuenta | La versión local no tiene cuotas: el plan «local» no limita documentos, páginas ni búsquedas, y admite ficheros de hasta 16 GB. Puede tener un solo usuario, varios sin cuenta externa o usar Clerk para iniciar sesión. ## Código abierto El código se publicará con licencia **EUPL-1.2** en [github.com/joseluissaorin/scholaris-v2](https://github.com/joseluissaorin/scholaris-v2). Mientras se termina de revisar, el repositorio es privado: no damos fecha. Las órdenes exactas de instalación estarán en su README, que manda sobre esta página. El SDK de Python ya tiene esa misma licencia. ## Qué se queda en tu máquina y qué no En la versión local, **tus ficheros, tu biblioteca, los vectores y el índice viven en tu disco**. Pero Scholaris no funciona del todo sin conexión: para leer páginas necesita al menos un lector en la nube. | Pieza | Sin conexión | Con un proveedor en la nube | | --- | --- | --- | | Leer páginas (escaneados, fotos, diapositivas) | No hay lector local todavía | Gemini, OpenRouter (Mistral OCR) o Workers AI | | Vectores | InferBox (Qwen3-VL Embedding, 2048 dimensiones) | Gemini Embedding 2 | | Reordenar y juzgar | InferBox | Jev (TypeSafe) o Workers AI | | Transcribir | InferBox | Gemini Transcribe o Whisper en Workers AI | | Redactar respuestas | InferBox | Gemini u OpenRouter | | Buscar y citar sobre lo ya leído | Sí: la búsqueda por palabras siempre; la de sentido, con InferBox | | La clave de Gemini es la única obligatoria en la práctica; las de OpenRouter y TypeSafe, Workers AI y OpenAlex son opcionales. Lo que sí funciona del todo sin conexión es **abrir y buscar un .spdf ya leído** con el SDK de Python (ver [El formato SPDF](https://scholaris.joseluissaorin.com/saber/spdf.md)). ## Lo mismo por fuera La versión local habla la misma API (v1 y v2) y el mismo MCP que la nube, así que el SDK de Python, los ejemplos de la [guía de la API](https://scholaris.joseluissaorin.com/api.md) y los agentes funcionan igual apuntando a `http://localhost:8790`. --- # Qué pasa con tus datos, proveedor a proveedor URL: https://scholaris.joseluissaorin.com/saber/privacidad > Dónde se guardan tus documentos, qué proveedor de IA lee qué, qué se manda a las bases bibliográficas abiertas, cómo se cifran tus claves y qué se borra cuando borras. ## Lo principal Tus documentos son tuyos. No entrenamos modelos con ellos. Para leerlos, transcribirlos y buscar en ellos, Scholaris manda páginas, audio y consultas a proveedores de IA, y esta hoja dice cuáles, para qué y qué se guarda en cada sitio. Si trabajas con material que no debe salir de tu ordenador, la [versión local](https://scholaris.joseluissaorin.com/saber/version-local.md) guarda todo en tu disco, aunque sigue necesitando un proveedor para leer páginas. No hay analítica de terceros ni anuncios en la web. Las únicas cookies son las de la sesión (Clerk) y una que recuerda que estás en la demostración. ## Dónde se guarda cada cosa Todo vive en Cloudflare, en la cuenta de Scholaris: | Qué | Dónde | | --- | --- | | Tu cuenta, tus claves de API (solo su huella), ajustes, uso, invitaciones y registro de accesos | D1, la base de datos de Cloudflare | | Los originales y las imágenes de las páginas | R2, el almacenamiento de objetos de Cloudflare, bajo una carpeta tuya | | Tu biblioteca leída (texto, anclas, fichas, índice) | Un Durable Object propio, con su base SQLite, solo para ti | | Los vectores | Cloudflare Vectorize, en un espacio de nombres propio | | La sesión | Clerk, que gestiona el inicio de sesión (correo y nombre) y los pagos | ## Qué lee cada proveedor | Proveedor | Qué recibe | Para qué | | --- | --- | --- | | Google Gemini (API de pago) | Imágenes de páginas, audio, la dirección de los vídeos de YouTube, pasajes y consultas | Leer páginas (Gemini 3.8 Flash y 3.5 Flash-Lite), transcribir (Gemini Transcribe), vectores (Gemini Embedding 2), redactar respuestas y extraer entidades | | Cloudflare Workers AI | Audio, imágenes de páginas fáciles y texto | Transcribir con Whisper, lector de reserva y del modo económico, vectores y reordenador de reserva | | OpenRouter (con Mistral OCR) | Páginas que los lectores anteriores no pudieron leer; texto para redactar si Gemini falla | Lector y redactor de reserva | | TypeSafe (Jev) | Pares de consulta y pasaje, afirmaciones y pasajes, candidatos de número de página | Reordenar resultados, juzgar citas y decidir folios dudosos | Estos proveedores reciben lo justo para cada tarea y se usan con sus API de pago o empresariales. Sus condiciones (no las nuestras) dicen cuánto tiempo conservan los datos y si los usan para algo más; por ejemplo, las condiciones de pago de la API de Gemini excluyen el uso de las peticiones para mejorar sus productos, aunque permiten guardarlas un tiempo limitado para detectar abusos. Si eso no te basta, usa la versión local con tus propias claves o con InferBox. ## Qué se manda a las bases bibliográficas abiertas Para completar y contrastar la ficha de cada documento, Scholaris pregunta a Crossref, OpenAlex, Open Library, Wikidata, Wikipedia, arXiv, DataCite y, como último recurso, Google Books. Lo que se manda es **el título, los autores, el ISBN o el DOI**, nunca el texto. Para enlazar entidades se manda a Wikidata el nombre de cada entidad. Las peticiones se identifican como «Scholaris/2» con la dirección de la web. ## Tus propias claves Puedes poner tus claves de Gemini, OpenRouter, TypeSafe, Mistral, Voyage, Cohere, Jina o ZeroEntropy para que tu biblioteca use tus cuentas. Se guardan cifradas con AES-256-GCM, con una clave derivada para cada usuario, y solo se usan si lo activas. ## Borrar - **Borrar un documento** borra también todo lo derivado: páginas, vectores, imágenes. - **Borrar el historial** de búsquedas, o pausar que se guarde. - **Exportar todo** en un ZIP antes de irte. - **Borrar la cuenta:** se borran tus claves, ajustes, uso, invitaciones, notificaciones y enlaces compartidos; se vacía tu biblioteca; los vectores y los ficheros se purgan en segundo plano. Queda una fila con tu identificador marcada como borrada, sin correo ni nombre, para que no se pueda reutilizar. La excepción: si alguien copió un documento de una biblioteca que le compartiste, su copia sigue siendo suya. ## Contacto Para cualquier cosa sobre tus datos, [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- # Planes, cupones y límites URL: https://scholaris.joseluissaorin.com/saber/planes > Un plan gratuito para empezar, Pro para trabajar de verdad y la versión local sin cuotas. Cuánto cabe en cada uno, cómo se cuentan las páginas y los minutos, y cómo funcionan los cupones. ## Los planes | | Gratis | Pro | Versión local | | --- | --- | --- | --- | | Documentos | 25 | 5000 | sin límite | | Páginas o minutos leídos al mes | 1500 | 60 000 | sin límite | | Búsquedas al día | 100 | 5000 | sin límite | | Autocitas al mes | 5 | 500 | sin límite | | Almacenamiento | 1 GB | 100 GB | tu disco | | Fichero más grande | 200 MB | 4 GB | 16 GB | | Peticiones a la API por minuto | 120 | 600 | sin límite práctico | El plan gratuito no caduca. El precio de Pro se ve en la aplicación (Ajustes), donde se contrata; los pagos los gestiona Clerk. La versión local es gratuita, pero pagas tú a los proveedores de IA que uses con tus claves. ## Cómo se cuenta - Una página de un PDF, de un EPUB o de una foto cuenta como una página; un minuto de audio o de vídeo, como una página. - Las cuotas mensuales se reinician el día 1 de cada mes y las diarias a medianoche, en hora UTC. - Volver a subir un fichero idéntico no cuenta: se detecta por su huella y se devuelve el que ya había. ## El modo económico Al llenar una biblioteca grande se puede elegir el modo económico: las páginas difíciles se leen por el lote de Gemini, a mitad de precio, y tardan horas en vez de segundos. En una comedia escaneada de 43 páginas, el coste bajó un 46 %. ## Cupones Un cupón es un código con la forma `SCHO-XXXX-XXXX` que concede un plan durante unos días o para siempre. Se canjea en Ajustes. Un cupón nunca baja de plan una cuenta; si ya tienes ese plan por tiempo limitado, lo alarga. Para evitar que se adivinen, se admiten diez intentos fallidos por hora, y Scholaris solo guarda una huella cifrada de cada código. Si crees que tu grupo o tu proyecto debería tener uno, escribe a [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). ## Probar sin cuenta La [demostración](https://scholaris.joseluissaorin.com/?demostracion) abre la aplicación con una biblioteca de ejemplo que vive en tu navegador: no se sube nada ni hace falta registrarse. --- # Lo que hemos medido, con sus condiciones URL: https://scholaris.joseluissaorin.com/saber/rendimiento > Tiempos y costes de lectura, exactitud del folio, calidad de búsqueda, citas inventadas y latencias, tal como salieron en el banco de pruebas del 6 de octubre de 2026, con lo que esas cifras no dicen. ## Las condiciones Todo se midió el **6 de octubre de 2026**, desde un Mac en Tenerife, por la red de casa, contra las API en la nube (Gemini 3.5 Flash-Lite y 3.8 Flash, Gemini Transcribe, Gemini Embedding 2 a 1536 dimensiones, Jev, Crossref y OpenAlex) y sin GPU local. **No se midió en Cloudflare**, que debería ser más estable; tampoco la latencia de «preguntar». Las cifras en dólares son el coste de los proveedores, no un precio. ## Leer un documento | Documento | Antes (primera Scholaris) | Ahora | Coste | | --- | --- | --- | --- | | Entrevista en vídeo, 54 min | 3 h 39 min | 45 s (30 s sin atribuir hablantes) | 0,41 $ | | Entrevista en vídeo (Cortázar), 2 h 2 min | las de 83 min tardaban 5 h 43 min | 70 s | 0,94 $ | | *The Discarded Image*, 245 p., con un OCR antiguo | sin registro | mediana de 85 s en 16 ejecuciones (de 37 a 175 s) | 0,64 $ | | *El casamiento en la muerte*, 43 p. escaneadas, siglo XVII | sin registro | 27 s | 0,20 $ | | *Attention Is All You Need*, 15 p. digitales | sin registro | 12 s | 0,04 $ | | *El perseguidor*, 37 p. digitales | sin registro | 15 s | 0,04 $ | Los tiempos son de principio a fin, con la conversión incluida. Las primeras páginas se pueden buscar antes: en menos de una décima de segundo en un PDF digital, en unos 3 a 6 segundos en un escaneado y en unos 3 a 6 segundos en un vídeo. La lentitud de algunas ejecuciones del libro de 245 páginas vino de la API o de la red, no del código. En los medios muy cortos (un minuto) el documento tarda ahora algo más en quedar listo que antes: de 5 a 6 segundos se pasó a entre 8 y 10. El **modo económico** bajó el coste del *Casamiento* de 0,206 $ a 0,111 $ (un 46 % menos), a cambio de que tardara 6,9 minutos en lugar de 19,7 segundos. Transcribir cuesta unos 0,005 $ por minuto. ## Leer bien - **Errores de carácter (CER)** en teatro del siglo XVII: 0,007 con Gemini 3.8 Flash, frente a 0,073 del lector de la primera Scholaris. Medido sobre una sola página transcrita a mano. - ***The Discarded Image***: la primera Scholaris dejaba 195 páginas vacías (74 000 caracteres en total); ahora salen 358 000 caracteres y 7 páginas vacías. - **Folio impreso:** 64 de 64 exactos en las páginas comprobadas a ojo (45 de 45 con el número a la vista). En *The Discarded Image*, 229 folios leídos coinciden con las etiquetas del PDF en 229 de 232 páginas. - **Hablantes:** atribución correcta en 20 de 21 turnos y en 23 de 23 en dos entrevistas (25 turnos al azar en cada una, anotados leyendo el texto). - **Fichas:** 55 de 55 comprobaciones correctas en 9 documentos tras contrastar con las fuentes externas (antes, 30 de 55). ## Buscar bien Sobre 183 consultas y 8879 juicios de relevancia. Los juicios son «de plata»: los hicieron dos modelos y un árbitro, no personas; coincidieron entre sí en un 82 % exacto y en un 99 % con un punto de margen, y de 30 revisados a mano, el autor estuvo de acuerdo con 29. | Sistema | nDCG@10 | Recall@20 | MRR | | --- | --- | --- | --- | | Primera Scholaris | 0,695 | 0,618 | 0,925 | | Solo léxica | 0,564 | | | | Solo visual | 0,199 | | | | Solo semántica | 0,804 | 0,860 | 0,931 | | Híbrida sin reordenar | 0,818 | 0,884 | 0,942 | | **La de ahora** (híbrida y reordenada) | **0,889** | **0,904** | **0,979** | Por tipo de consulta (nDCG@10 del sistema actual): entre lenguas 0,897; literales 0,880; audio y vídeo 0,882; grafía antigua 0,814; visuales 0,950 (solo 4 consultas). **Grafía antigua:** en *El casamiento en la muerte* de Lope, con 30 consultas, la exhaustividad media por consulta pasó del 32 % al 100 % con la capa de grafía modernizada; en documentos modernos, 250 de 250 búsquedas devolvieron exactamente lo mismo. **Latencia:** mediana de unos 650 ms desde casa, sin cachés (unos 360 ms el vector de la consulta y unos 300 ms el juez); sin reordenar, unos 350 ms con nDCG 0,818. En las ejecuciones medidas, la mediana y el percentil 95 fueron 687 y 856 ms. ## Citar bien Sobre 41 afirmaciones (30 respaldadas y 11 que no lo estaban): **precisión del 95,1 %, exhaustividad del 96,7 %, cero citas inventadas**, cero citas en afirmaciones sin respaldo y el veredicto de verificación, correcto en el 100 %. Cada afirmación tardó 998 ms y costó 0,0054 $. ## Lo que estas cifras no dicen - Son de una red doméstica, no de Cloudflare, y no incluyen la latencia de «preguntar». - Quedan 1594 juicios de relevancia con árbitro o discrepancia pendientes de revisión humana. - El CER se calculó sobre una sola página. - El banco es pequeño y lo hizo quien hace Scholaris. Si quieres repetirlo con tus documentos, escríbenos. --- # Preguntas frecuentes URL: https://scholaris.joseluissaorin.com/saber/preguntas > Si inventa citas, si sirve con libros escaneados del siglo XVII o con entrevistas en vídeo, qué pasa con tus datos, si funciona sin conexión y otras dudas, contestadas sin rodeos. ## Sobre las citas ### ¿Inventa citas? No. Ninguna cita la escribe un modelo: el autor, el año y la página salen de la ficha del documento y del ancla que se guardó al leerlo. Cuando un modelo redacta una respuesta, solo puede citar los pasajes que se le dan, y cada nota se comprueba contra el texto antes de enseñártela. Si no hay pasaje que respalde algo, la respuesta lo dice. En el banco de 41 afirmaciones, las citas inventadas fueron cero. Ver [Citas](https://scholaris.joseluissaorin.com/saber/citas.md). ### ¿La página es la del PDF o la del libro? La del libro: el número impreso en el papel, aunque el PDF empiece en la cubierta, el libro use romanos en el prólogo o se salte láminas. Si el libro no imprime número, se cita la posición física entre corchetes, «p. [12]», en lugar de inventar uno. Ver [Formatos y anclas](https://scholaris.joseluissaorin.com/saber/formatos.md). ### ¿Qué estilos de cita tiene? APA, Chicago (autor-fecha y notas), MLA, Harvard, ISO 690 (autor-fecha en castellano y en inglés, y numérico) e IEEE, en castellano, inglés, francés e italiano. Exporta BibTeX, RIS y CSL-JSON, e inserta las citas en un .docx con notas al pie de Word de verdad. ## Sobre lo que lee ### ¿Funciona con libros escaneados del siglo XVII? Sí, y es uno de los casos para los que está hecho. Lee las páginas con un modelo de visión que respeta la ortografía original («dixo», «muger», u/v, ç) y deja fuera titulillos, reclamos y signaturas; reconoce la foliación por hojas («23r», «23v»). Para que esas grafías aparezcan cuando buscas «dijo» o «mujer», cada pasaje lleva una capa de grafía modernizada solo para buscar. En una comedia de Lope de 43 páginas, los errores de carácter fueron un 0,7 % y la exhaustividad de búsqueda pasó del 32 % al 100 %. Ver [Búsqueda](https://scholaris.joseluissaorin.com/saber/busqueda.md#grafia). ### ¿Y con entrevistas en vídeo? También. Transcribe palabra a palabra, distingue quién habla y le pone nombre, parte el texto en tramos de 30 a 60 segundos y cita el minuto exacto («12:04-12:40»). Los fotogramas también se buscan. Una entrevista de dos horas quedó lista en 70 segundos. YouTube se lee desde la dirección, sin descargar el vídeo. Ver [Audio y vídeo](https://scholaris.joseluissaorin.com/saber/reproductor.md). ### ¿Lee latín? Sí. La capa de búsqueda iguala u/v, i/j, æ/œ y los enclíticos, y reduce las palabras a su raíz con el lematizador de Schinke y otros, así que «urbis» encuentra «Vrbs». ### ¿Qué formatos acepta? PDF (digitales y escaneados), fotos (JPG, PNG, HEIC…), EPUB, Word, ODT, RTF, HTML, Markdown, PowerPoint, Keynote, ODP, Excel, ODS, CSV, audio (MP3, WAV, M4A…), vídeo (MP4, MOV, WebM…), webs, YouTube, Vimeo y pódcast. La lista completa, con el ancla de cada uno, está en [Formatos y anclas](https://scholaris.joseluissaorin.com/saber/formatos.md). ## Sobre tus datos ### ¿Qué pasa con mis datos? Se guardan en Cloudflare, en un espacio propio para tu cuenta, y no se usan para entrenar modelos. Para leerlos y buscar en ellos se mandan páginas, audio y consultas a Google Gemini, Cloudflare Workers AI, OpenRouter y TypeSafe, cada uno para una tarea concreta; a las bases bibliográficas abiertas solo se mandan títulos, autores, ISBN y DOI. Puedes exportarlo todo y borrar tu cuenta cuando quieras. El detalle, proveedor a proveedor, está en [Privacidad](https://scholaris.joseluissaorin.com/saber/privacidad.md). ### ¿Alguien más puede ver mi biblioteca? Solo si la compartes: invitando a alguien por correo o creando un enlace público, que puedes proteger con contraseña, hacer caducar o revocar. Las páginas de esos enlaces piden a los buscadores que no las indexen. ### ¿Puedo usarlo sin conexión? En parte. La versión local guarda todo en tu ordenador y, con InferBox, hace allí los vectores, la transcripción y las respuestas; pero para leer páginas necesita todavía un proveedor en la nube. Lo que sí funciona del todo sin conexión es abrir y buscar un .spdf ya leído con el SDK de Python. Ver [Versión local](https://scholaris.joseluissaorin.com/saber/version-local.md). ### ¿Es de código abierto? El SDK de Python lo es (EUPL-1.2). El resto del código se publicará con la misma licencia cuando termine su revisión; aún no hay fecha. ## Sobre el uso ### ¿Escribe mi trabajo? No, y no lo hará. Busca, señala, cita y verifica; pensar y escribir siguen siendo tuyos. La función «citar» devuelve tu propio texto con las citas puestas, no un texto nuevo. ### ¿Busca artículos en internet? No: trabaja sobre lo que tú le das. Para descubrir literatura que aún no tienes, usa un buscador académico y trae aquí lo que encuentres. El grafo de citas sí te dice qué obras citan tus libros y no tienes. ### ¿Cuánto cuesta? Hay un plan gratuito (25 documentos, 1500 páginas o minutos al mes, 100 búsquedas al día) y un plan Pro, cuyo precio se ve en la aplicación. La versión local no tiene cuotas. Ver [Planes](https://scholaris.joseluissaorin.com/saber/planes.md). ### ¿Lo puede usar un agente o un programa? Sí, con la API v1 o con el servidor MCP (Claude, Cursor y otros). Ver [La API v1 y el servidor MCP](https://scholaris.joseluissaorin.com/saber/api-y-mcp.md) y la [hoja para agentes](https://scholaris.joseluissaorin.com/agentes.md). ### ¿Quién lo hace? José Luis Saorín Ferrer, filólogo y programador, en Santa Cruz de Tenerife. Escríbele a [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- # Glosario de Scholaris URL: https://scholaris.joseluissaorin.com/saber/glosario > Las palabras que usa Scholaris, de folio y ancla a pliego, imprenta, SPDF o vigilante, con lo que significan aquí y, cuando viene a cuento, en la tradición del libro. ## Las palabras Ancla : El sitio exacto de un pasaje dentro de su documento: la página física y el folio impreso, el segundo de un audio, la diapositiva, la hoja y las filas de una tabla, o la sección y el párrafo de una web con su fecha de consulta. Toda cita se escribe desde un ancla. Autocita : La función que lee tu borrador, busca en tu biblioteca un pasaje para cada afirmación y propone la cita con su página; tú aceptas o descartas. Cuaderno : Notas en Markdown con fichas de pasajes verificados que se pueden volver a comprobar contra la biblioteca. Espacio vectorial : El modelo, la versión y las dimensiones con que se calcularon unos vectores, por ejemplo «gemini-embedding-2@1536». Un .spdf puede llevar varios. Ficha : Los metadatos de un documento (autores, título, año, edición, editorial, DOI, ISBN…), cada campo con su procedencia. Folio : En el libro antiguo, cada hoja numerada, con su recto (r) y su vuelto (v): «fol. 23v». En Scholaris, por extensión, el número de página impreso en el papel, frente a la posición de la página en el fichero. Foliación : La numeración de un libro: por páginas, por hojas, en romanos y arábigos, con saltos y láminas sin numerar. Scholaris la reconstruye página a página. Fragmento : El pasaje que se busca y se cita: un trozo de texto con su ancla de inicio y de fin. Grafía modernizada : Una copia del texto con la ortografía puesta al día («dixo» → «dijo»), guardada aparte y usada solo para buscar. El texto que se cita conserva siempre la grafía original. Huella : El resumen SHA-256 de un fichero; dos ficheros con la misma huella son el mismo fichero. Imprenta : El conversor de Scholaris: reconoce el tipo de cada fichero y lo prepara para leerlo (páginas como imágenes, texto, audio, fotogramas). Trabaja en el navegador y, para lo que el navegador no sabe hacer, en un contenedor del servidor. Jev : El modelo de TypeSafe que hace de juez: reordena resultados, decide si un pasaje respalda una afirmación y elige entre candidatos de número de página. Lector : El modelo que lee páginas como imágenes y devuelve su texto fiel, sus notas, su cabecera y su pie, y el folio que ve impreso. Manícula : La manita con el índice extendido que los lectores antiguos dibujaban en los márgenes para señalar un pasaje. Es el emblema de Scholaris. Pliego : En imprenta, la hoja grande que, doblada, forma un cuadernillo del libro. En Scholaris, el grupo de cuatro páginas que se manda de una vez al lector. Procedencia : De dónde sale cada dato: qué fuente dio la ficha, qué lector leyó la página, con qué confianza y quién lo corrigió. Reclamo : En el libro antiguo, la palabra que se imprime al pie de una página y repite la primera de la siguiente. Scholaris la deja fuera del texto del pasaje. SPDF : El formato abierto de Scholaris para un documento ya leído: una base de datos SQLite comprimida con gzip, con texto, anclas, figuras, vectores y procedencia. Versión actual: 4.1. Paquete .scholaris : Una biblioteca entera en un ZIP: un .spdf por documento, un manifiesto y un aviso de derechos. Tramo : En audio y vídeo, la unidad citable: entre 30 y 60 segundos de transcripción que acaban en final de frase. Unidad : Lo que se cita como un todo: una página, un tramo de audio, una diapositiva, un trozo de hoja de cálculo. Vigilante : Una búsqueda guardada que se repite sola y avisa cuando cambia lo que encuentra. --- # Scholaris al lado de Zotero, Elicit, NotebookLM y los demás URL: https://scholaris.joseluissaorin.com/saber/alternativas > Qué hace cada herramienta, cuándo conviene otra antes que Scholaris y qué hace Scholaris que las demás no se proponen. Sin cifras ajenas que no podamos comprobar. ## Cómo leer esta hoja No hemos medido las otras herramientas con nuestro banco de pruebas, así que aquí no hay cifras suyas ni tablas de casillas marcadas. Lo que sigue describe para qué está pensada cada una, a la fecha de esta revisión, y en qué se diferencia Scholaris. Las herramientas cambian deprisa: si algo de lo que decimos de ellas ha dejado de ser cierto, avísanos y lo corregimos. ## Zotero Un gestor de referencias libre y de código abierto: recoge referencias desde el navegador, guarda los PDF, los anota y da formato a las citas en Word, LibreOffice y Google Docs con estilos CSL. Para coleccionar y ordenar referencias es difícil de mejorar, y Scholaris no intenta sustituirlo. **Cuándo elegir Zotero:** para construir y mantener tu bibliografía, colaborar en grupos de referencias y citar mientras escribes en el procesador de textos. **Lo que añade Scholaris:** leer el contenido de esas fuentes (también escaneados antiguos, audio y vídeo), encontrar el pasaje por su sentido y devolver la cita con la página impresa exacta. Se llevan bien: Scholaris importa el BibTeX o el RIS de Zotero con su carpeta de ficheros y casa cada entrada con su documento. ## Elicit Un asistente de investigación que busca en la literatura científica publicada, resume artículos y extrae datos de ellos en tablas para revisiones sistemáticas. **Cuándo elegir Elicit:** para descubrir artículos que todavía no tienes y comparar muchos estudios empíricos a la vez. **Lo que añade Scholaris:** trabaja sobre tu propia biblioteca, no sobre un corpus de artículos: libros escaneados del siglo XVII, ediciones concretas, entrevistas grabadas, apuntes; y cita la página impresa o el segundo de tu ejemplar. ## NotebookLM La herramienta de Google para conversar con las fuentes que subes, con referencias a los pasajes de origen y resúmenes en audio. **Cuándo elegir NotebookLM:** para hacerte una idea rápida de un conjunto pequeño de fuentes, gratis y sin instalar nada, si te parece bien que las procese Google dentro de su producto. **Lo que añade Scholaris:** citas en estilos bibliográficos (APA, Chicago, MLA, ISO 690…) con el folio impreso, listas para una nota al pie; foliación de libros antiguos y capa de grafía modernizada; verificación con lógica temporal; un formato abierto (SPDF) para llevarte la biblioteca ya leída; API y MCP, y una versión que se instala en tu ordenador. Y no te resume los libros: no es para eso. ## ChatGPT o Claude con ficheros subidos Los asistentes generales leen los ficheros que adjuntas a una conversación y contestan sobre ellos, a menudo muy bien. **Cuándo elegirlos:** para conversar, explorar ideas o redactar con ayuda, sobre unos pocos documentos. **Lo que añade Scholaris:** una biblioteca persistente de miles de documentos en lugar de los ficheros de una conversación; citas que salen del ancla guardada y no del modelo, comprobadas antes de llegar a ti; la página impresa y el segundo; y la regla de no escribir por ti. De hecho, se pueden usar juntos: Claude puede consultar tu biblioteca de Scholaris por [MCP](https://scholaris.joseluissaorin.com/saber/api-y-mcp.md) y citar desde ella. ## Adobe Acrobat con IA El asistente de Acrobat responde preguntas sobre PDF y genera resúmenes dentro del lector de PDF más extendido. **Cuándo elegirlo:** si ya trabajas en Acrobat y te basta con preguntar sobre los PDF que tienes abiertos. **Lo que añade Scholaris:** muchos formatos además del PDF (EPUB, audio, vídeo, webs, presentaciones), búsqueda en toda la biblioteca a la vez, estilos de cita y página impresa, y el texto de los escaneados antiguos leído con fidelidad. ## En resumen | Si necesitas… | Mira… | | --- | --- | | Coleccionar referencias y citar mientras escribes en Word | Zotero (y, para el contenido, Scholaris al lado) | | Descubrir artículos científicos que aún no tienes | Elicit o un buscador académico | | Hacerte una idea rápida de unas pocas fuentes | NotebookLM o un asistente general | | Encontrar en tu biblioteca el pasaje exacto y citarlo con su página o su segundo | Scholaris | --- # Registro de cambios URL: https://scholaris.joseluissaorin.com/saber/cambios > Lo que ha cambiado en Scholaris, de la primera versión a la segunda, contado por hitos y en el orden en que se hicieron. ## Octubre de 2026 - **6 de octubre:** scholaris.joseluissaorin.com pasa a servir la segunda versión de Scholaris, en Cloudflare. La primera versión queda en reserva. - Esta base de conocimiento, con un gemelo en Markdown de cada hoja, [/llms.txt](https://scholaris.joseluissaorin.com/llms.txt), [/llms-full.txt](https://scholaris.joseluissaorin.com/llms-full.txt) y la [hoja para agentes](https://scholaris.joseluissaorin.com/agentes.md). - El SDK de Python, listo para publicarse en PyPI como `scholaris-sdk`, con licencia EUPL-1.2. - Exportar un .spdf desde el servidor, con un tope de 24 MB de ficheros incrustados y aviso de lo que queda fuera. - Rehacer la ficha de un documento cuando las pruebas de su identidad cambian (otro episodio, otro invitado), con modo simulado. - Rehacer solo las figuras de un documento, con el coste estimado antes. ## La segunda versión, en el orden en que se hizo La segunda versión de Scholaris se escribió de nuevo. Sus hitos, del primero al último: 1. **Cimientos:** el formato SPDF 4.0 y la imprenta, el conversor que lee cualquier fichero en el navegador. 2. **Búsqueda** por tres caminos fundidos (léxico, semántico y visual) y respuestas con citas validadas; los estilos CSL oficiales. 3. **Lectura por lotes de páginas**, folios impresos, fichas, fragmentos, vectores y figuras; el mapa de conceptos. 4. **Citas:** lógica temporal, autocita con juez, inserción en .docx y BibTeX. 5. **La API en Cloudflare Workers**, con inicio de sesión de Clerk y claves propias cifradas; el SPDF en Workers y la migración de los SPDF antiguos. 6. **La aplicación web** (biblioteca, lector, búsqueda, escritura, exploración), la **versión local** sobre Node y SQLite, el **ejecutable de escritorio** y la imagen de Docker. 7. **Un lector nuevo**, Gemini 3.8 Flash, elegido con el banco de pruebas. 8. **Medios:** YouTube por su dirección, pódcast, Vimeo y enlaces directos. 9. **El servidor MCP con OAuth 2.1** y la migración de las cuentas de la primera versión. 10. **La lectura en tres tiempos** (legible, buscable, con vectores), el **modo económico** y el **grafo de entidades** enlazado con Wikidata. 11. **El banco de calidad:** nDCG@10 de 0,889 y ninguna cita inventada. 12. **La búsqueda en dos tiempos** y el **reproductor nuevo**, con la palabra que suena resaltada. 13. **La API v1** con su guía en [/api](https://scholaris.joseluissaorin.com/api.md), la versión local para varios usuarios y las bibliotecas compartidas. 14. **Los dibujos y el movimiento** de la aplicación; fichas de programas de radio y televisión; OpenAlex con caché. 15. **La puesta a punto para producción.** ## La primera versión La primera Scholaris corría en un servidor doméstico y guardaba los documentos en SPDF 1 a 3. Sus SPDF se siguen abriendo: se migran solos al formato 4. Comparada con ella, la segunda lee una entrevista de 54 minutos en 45 segundos en lugar de tres horas y 39 minutos, y encuentra mejor (nDCG@10 de 0,695 a 0,889). Ver [Rendimiento](https://scholaris.joseluissaorin.com/saber/rendimiento.md). --- # Para agentes: cómo leer Scholaris y cómo usarlo en nombre de alguien URL: https://scholaris.joseluissaorin.com/agentes > Qué puede leer un agente en esta web y en qué formato, cómo actuar sobre la biblioteca de una persona con la API v1 o el MCP, con qué permisos y límites, y las reglas para citar sin inventar. ## Si solo quieres entender qué es Scholaris Todo lo público está pensado para leerse sin JavaScript y sin raspar HTML. Cada página tiene un gemelo en Markdown en la misma dirección terminada en `.md`, y también se sirve en Markdown si lo pides con la cabecera `Accept: text/markdown`. Las páginas llevan además JSON-LD de schema.org (SoftwareApplication, TechArticle, FAQPage, HowTo, DefinedTermSet). ```sh # Cualquier página pública, en Markdown curl -s https://scholaris.joseluissaorin.com/saber/formatos.md curl -s -H "Accept: text/markdown" https://scholaris.joseluissaorin.com/saber/formatos # Todo de una vez curl -s https://scholaris.joseluissaorin.com/llms.txt curl -s https://scholaris.joseluissaorin.com/llms-full.txt ``` ## Los ficheros para máquinas | Dirección | Qué es | | --- | --- | | [/llms.txt](https://scholaris.joseluissaorin.com/llms.txt) | Índice breve con un enlace a cada página en Markdown | | [/llms-full.txt](https://scholaris.joseluissaorin.com/llms-full.txt) | El texto entero de todas las páginas públicas, en castellano y en inglés | | /cualquier/pagina.md | El gemelo en Markdown de cada página pública | | [/api/v1/llms.txt](https://scholaris.joseluissaorin.com/api/v1/llms.txt) | Cómo usar la API sin inventar citas | | [/api/v1/openapi.json](https://scholaris.joseluissaorin.com/api/v1/openapi.json) | Especificación OpenAPI 3.1 de la API v1 | | [/.well-known/api-catalog](https://scholaris.joseluissaorin.com/.well-known/api-catalog) | Catálogo de API (RFC 9727) | | [/.well-known/mcp/server-card.json](https://scholaris.joseluissaorin.com/.well-known/mcp/server-card.json) | Tarjeta del servidor MCP | | [/.well-known/oauth-protected-resource/mcp](https://scholaris.joseluissaorin.com/.well-known/oauth-protected-resource/mcp) | Metadatos OAuth del recurso MCP | | [/.well-known/oauth-authorization-server](https://scholaris.joseluissaorin.com/.well-known/oauth-authorization-server) | Metadatos del servidor de autorización OAuth | | [/.well-known/security.txt](https://scholaris.joseluissaorin.com/.well-known/security.txt) | Dónde avisar de un problema de seguridad | | [/sitemap.xml](https://scholaris.joseluissaorin.com/sitemap.xml) | Todas las páginas públicas, con su fecha y sus lenguas | | [/robots.txt](https://scholaris.joseluissaorin.com/robots.txt) | Qué se puede rastrear | ## Qué se puede rastrear Las páginas públicas (la portada, esta base de conocimiento, la guía de la API y esta hoja) se pueden leer, indexar, citar y usar para responder, también para entrenar modelos: está dicho en [/robots.txt](https://scholaris.joseluissaorin.com/robots.txt), con una línea para cada rastreador conocido (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot y otros). **No se rastrea** la aplicación, la API ni las bibliotecas de nadie: los enlaces públicos que comparten las personas piden expresamente no ser indexados, y su contenido es suyo. ## Si actúas en nombre de una persona Necesitas que esa persona te dé acceso; no hay acceso anónimo a ninguna biblioteca. Hay dos caminos: 1. **MCP con OAuth.** Si tu cliente habla MCP, conecta con `https://scholaris.joseluissaorin.com/mcp`. Te registras solo (registro dinámico de clientes), la persona entra en Scholaris y te concede acceso de solo lectura. Herramientas: `search`, `cite`, `open_page` y `verify_claim`. 2. **API v1 con una clave.** La persona crea una clave en Ajustes → Claves de API y te la da. Con el alcance `lectura` puedes buscar, leer, preguntar y verificar; con `escritura`, además subir, borrar y citar un texto entero; con `mcp`, usarla en el servidor MCP. Base: `https://scholaris.joseluissaorin.com/api/v1`. ```sh curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=la música me metía en el tiempo" -d k=5 -d formato=markdown ``` Todo está explicado en [La API v1 y el servidor MCP](https://scholaris.joseluissaorin.com/saber/api-y-mcp.md) y, con una sesión grabada, en la [guía de la API](https://scholaris.joseluissaorin.com/api.md). ## Límites que tienes que respetar - **Ritmo:** 120 peticiones por minuto en el plan gratuito y 600 en Pro. Si recibes un 429, espera los segundos de `Retry-After`. - **Cuotas** de la persona: búsquedas al día, páginas o minutos al mes, autocitas al mes (ver [Planes](https://scholaris.joseluissaorin.com/saber/planes.md)). Un 402 quiere decir que se han agotado: díselo, no insistas. - **Ficheros:** hasta 95 MB por petición en la API v1. - **Reintentos:** manda `Idempotency-Key` al subir y al citar para no duplicar nada. - **Lo que tarda:** subir y citar esperan por defecto; para no bloquearte, `esperar=0` o `Prefer: respond-async`, y consulta `progreso_url`. ## Las reglas para citar 1. Cita solo lo que devuelve Scholaris, y copia `cita` y `localizador` tal cual. 2. Si citas literalmente, copia el `texto` literal; si parafraseas, pon igualmente la cita. 3. Da el `enlace`: es como la persona comprueba la página o el segundo. 4. Si la búsqueda no encuentra nada, dilo. No rellenes el hueco de memoria ni inventes una página. 5. No presentes una respuesta de `preguntar` como si fuera tuya: sus notas al pie son la parte importante. 6. No escribas el trabajo de la persona haciéndolo pasar por suyo. Scholaris existe para que encuentre y piense ella. ## Errores Todas las respuestas de error tienen un `codigo` estable, un `mensaje` en castellano, un `message` en inglés y un enlace a la documentación. Los campos de la API están en castellano y en minúsculas. ## Contacto Para avisar de un fallo, pedir más ritmo o proponer una integración: [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- URL: https://scholaris.joseluissaorin.com/acerca # Una mano que señala la página y luego se aparta. > Scholaris lee tus libros, tus artículos, tus apuntes y tus entrevistas, y cuando le preguntas algo no te devuelve una opinión: te señala el pasaje, la página impresa o el minuto exacto en que se dijo, para que el pensamiento lo sigas poniendo tú. Esta es la portada de Scholaris, escrita como un ensayo en ocho capítulos. Los detalles técnicos (formatos, citas, privacidad, API, rendimiento) están en la base de conocimiento: https://scholaris.joseluissaorin.com/saber.md ## Capítulo primero, que trata de la abundancia y de lo poco que en ella se encuentra En algún disco duro, dentro de una carpeta que se llamaba «leer» y que luego se llamó «leer de verdad» y después «tesis_final_2», está el párrafo que necesitas. Lo subrayaste una tarde. Sabes que estaba en una página par, más o menos hacia el final de un capítulo, y que al lado había una nota tuya a lápiz que decía algo parecido a «¡sí!». No lo encuentras. Nunca hubo tanto que leer ni fue tan fácil guardarlo, y sin embargo la experiencia más común de quien estudia sigue siendo esa: saber que algo existe y no poder llegar hasta ello. Tenemos bibliotecas que caben en un bolsillo y que se comportan como aquella otra, interminable, de galerías hexagonales, donde están todos los libros posibles y ninguno se deja abrir por la página justa. La abundancia no ha resuelto el problema de encontrar; lo ha agrandado hasta volverlo invisible, porque ahora la culpa parece nuestra. > Cf. J. L. Borges, «La biblioteca de Babel», en *Ficciones* (1944). Sin número de página: esta portada tampoco cita lo que no ha comprobado. ## Capítulo segundo, que trata del archivo y de quién decide lo que se puede decir Un archivo nunca es inocente. Lo que se guarda, el orden en que se guarda y lo que se deja fuera deciden, antes que nadie, lo que luego se podrá pensar; durante siglos lo decidieron los copistas, los bibliotecarios y los índices, y hoy lo deciden, cada vez más, unas máquinas que no enseñan sus estanterías. **Cuando no sabes de dónde sale lo que lees, no puedes discutirlo, y lo que no se puede discutir acaba mandando.** Por eso nos importa tanto la procedencia, que no es un adorno académico ni una nota al pie que nadie mira, sino lo mínimo que necesita una lectura para ser tuya: saber quién lo dijo, dónde, en qué edición, en qué página y con qué grado de certeza, y poder ir a comprobarlo. > Cf. M. Foucault, *La arqueología del saber* (1969), «El *a priori* histórico y el archivo». ## Capítulo tercero, que trata de las máquinas que contestan El buscador te devuelve diez enlaces y una página de anuncios; el asistente conversacional, una respuesta redonda, segura de sí misma, sin un solo folio. El primero te deja en la puerta de un edificio ajeno; el segundo te cuenta lo que hay dentro sin dejarte entrar y, a veces, se inventa los libros, los autores y las páginas con la misma voz tranquila con la que acierta. Los dos aplanan. Convierten la diferencia entre un texto y otro —entre una edición de 1925 y una traducción de 1970, entre lo que alguien dijo y lo que se dice que dijo— en una superficie lisa donde todo pesa lo mismo. Y los dos, sin quererlo, te quitan el trabajo que más te importaba hacer: leer el pasaje, desconfiar de él, ponerlo al lado de otro y pensar. ## Capítulo cuarto, que trata de la mano que señala En los márgenes de los libros antiguos aparece una y otra vez una manita con el índice extendido. Los paleógrafos la llaman manícula, y servía exactamente para eso: para decirle a un lector futuro «mira aquí», sin decirle qué tenía que pensar de lo que iba a ver. Scholaris quiere ser esa mano. Cuando le preguntas algo, te devuelve el pasaje, el libro y la página impresa —la de verdad, la que pondrías en una nota al pie— o el minuto exacto de la entrevista en que alguien lo dijo. Si no lo encuentra, lo dice. **No inventa citas:** cada una se comprueba contra el texto que tú le diste antes de enseñártela, y cada dato lleva su procedencia a la vista, de dónde salió, con qué confianza y quién lo corrigió. > «En un lugar de la Mancha, de cuyo nombre no quiero acordarme, no ha mucho tiempo que vivía un hidalgo de los de lanza en astillero…» > Miguel de Cervantes, *El ingenioso hidalgo don Quijote de la Mancha*, Madrid, Juan de la Cuesta, 1605, fol. 1r Y si lo que tienes es una grabación, el minuto 12:04, palabra a palabra. ## Capítulo quinto, que trata de cómo una biblioteca ordenada empieza a pensar contigo Una biblioteca bien leída deja ver relaciones que nadie había anotado: dos autores que no se citan y hablan de lo mismo, un concepto que cambia de nombre al cruzar una frontera, una pregunta que vuelve cada cuarenta años con otra ropa. Scholaris lee toda tu biblioteca a la vez y dibuja esas relaciones (personas, obras, lugares, conceptos, fechas) como un mapa del cielo, con constelaciones que no estaban en ningún libro y que solo aparecen cuando los libros se ponen juntos. Las líneas las traza la máquina, siempre ancladas a la página de la que salen; qué significan, cuáles importan y adónde llevan, eso lo decides tú. Queremos una estructura que multiplique el pensamiento en lugar de cerrarlo, que funcione menos como un árbol, con su tronco y su jerarquía, y más como un rizoma, donde cualquier punto puede conectarse con cualquier otro y cada conexión es el principio de una idea que todavía no existe. > Cf. G. Deleuze y F. Guattari, *Mil mesetas* (1980), «Introducción: rizoma». ## Capítulo sexto, que trata de lo que Scholaris no hará nunca No escribirá tu ensayo. No resumirá un libro para que no tengas que leerlo. No decidirá por ti qué cita es la buena ni qué autor tiene razón. Pensar, juzgar y escribir siguen siendo tuyos, y lo son por una cuestión de principio: un pensamiento que no ha pasado por el cuerpo de quien lo firma se parece demasiado a un rumor. ## Capítulo séptimo, que trata de lo que hace, contado deprisa - **Cualquier cosa entra** (PDF · EPUB · DOCX · MP3 · MP4). PDF digitales o escaneados, fotos de páginas, EPUB, Word, presentaciones, audio, vídeo y enlaces: todo se vuelve una biblioteca que se puede buscar y citar. - **La página impresa** (p. 145 · xiv · fol. 1r). El folio de verdad, el que pondrías en la nota, aunque el libro empiece en romanos o la paginación salte. - **El minuto exacto** (12:04). Entrevistas, clases y grabaciones transcritas palabra a palabra; cada cita salta a su segundo. - **Preguntas a toda la biblioteca** (citas verificadas). En cualquier lengua, también en castellano antiguo o en latín, con respuestas cuyas citas se comprueban una a una antes de llegar a ti. - **Tus citas, en tu estilo** (CSL · BibTeX · RIS). APA, MLA, Chicago o el que pida tu revista; BibTeX y RIS; y una autocita que repasa tu borrador y propone cada referencia con su página. - **En tu ordenador o en la nube** (local · nube). Se instala en casa, sin cuenta y con tu biblioteca en tu disco (la lectura usa tus propias claves de IA o tu propio servidor de inferencia), o se usa en la nube desde cualquier sitio. - **Un formato abierto y tuyo** (.spdf). Tu biblioteca entera, ya leída, en un archivo SPDF abierto que te llevas cuando quieras. ## Capítulo octavo y último, que trata de para quién es Para quien lee para escribir. Para la doctoranda con trescientos PDF y una directora que pide la página; para el profesor que prepara clase con veinte años de subrayados; para la periodista que vuelve a una entrevista de hace dos años buscando una frase; para el traductor, la archivera, el estudiante de primero que todavía no sabe que va a necesitarlo, y para cualquiera que haya dicho alguna vez «lo leí en algún sitio». ## Aquí se acaba la portada Scholaris lo hace José Luis Saorín Ferrer, filólogo, programador y artista, en Santa Cruz de Tenerife. Los dibujos de esta página están escritos trazo a trazo, como quien calca, y un programa les pone el pulso: tiemblan siempre igual, pero tiemblan. *Santa Cruz de Tenerife, otoño de 2026* Empezar: https://scholaris.joseluissaorin.com/?entrar · Probar sin cuenta: https://scholaris.joseluissaorin.com/?demostracion --- URL: https://scholaris.joseluissaorin.com/api # Subes cualquier cosa, buscas, preguntas y citas. > Una cabecera, nueve verbos y JSON de ida y vuelta (o Markdown, que los agentes leen mejor). Cada pasaje llega con su página impresa o su segundo exactos y un enlace al lector, y ninguna cita sale de un modelo: todas salen del ancla guardada. Base: `https://scholaris.joseluissaorin.com/api/v1` · OpenAPI: https://scholaris.joseluissaorin.com/api/v1/openapi.json · llms.txt: https://scholaris.joseluissaorin.com/api/v1/llms.txt ## Empieza en un minuto 1. Crea una clave en Ajustes → Claves de API, con los alcances de lectura y escritura. Se enseña una sola vez: guárdala. 2. Sube un fichero. La petición espera a que esté leído (hasta un minuto; si tarda más, responde 202 con dónde preguntar). 3. Busca. Cada pasaje trae su cita lista para pegar, el localizador exacto y el enlace al lector en esa página. ```sh export SCHOLARIS=sch_… # tu clave curl -s https://scholaris.joseluissaorin.com/api/v1/documentos -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/pdf" --data-binary @articulo.pdf curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=atención escalada" -d k=3 -d formato=markdown ``` Los ejemplos se pueden copiar y pegar tal cual: solo hace falta la variable SCHOLARIS con tu clave. ## Los nueve verbos Todo cuelga de una misma dirección base. Los nombres de los campos están en castellano y en minúsculas. ### POST /api/v1/documentos Sube cualquier cosa: el fichero como cuerpo (PDF, EPUB, DOCX, presentaciones, hojas, audio, vídeo, imágenes), un formulario multipart con el campo `archivo` o un JSON con `url` (una web, un PDF, YouTube, Vimeo, un pódcast). Si el fichero ya estaba (misma huella SHA-256), devuelve el que hay con `duplicado`. ```sh # Un fichero, como cuerpo (con su tipo y su nombre) curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos?nombre=libro.epub" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/epub+zip" --data-binary @libro.epub # Una URL (web, PDF, YouTube, Vimeo, pódcast), sin esperar curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos?esperar=0" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/json" -d '{"url": "https://www.youtube.com/watch?v=fNk_zzaMoSs"}' # Un formulario, con la ficha curl -s https://scholaris.joseluissaorin.com/api/v1/documentos -H "Authorization: Bearer $SCHOLARIS" \ -F archivo=@entrevista.mp3 -F titulo="A fondo" -F autores="Cortázar, Julio" -F anio=1977 ``` ### GET /api/v1/documentos Lista la biblioteca. Filtra con `q` (título o autor) y `estado`; pagina con `cursor`. ```sh curl -sG https://scholaris.joseluissaorin.com/api/v1/documentos -H "Authorization: Bearer $SCHOLARIS" -d estado=listo -d formato=markdown ``` ### GET /api/v1/documentos/{id} La ficha: título, autores, año, estado, progreso mientras se procesa y la referencia completa cuando está listo. Con `esperar=30` espera a que termine. ```sh curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos/ID?esperar=30" -H "Authorization: Bearer $SCHOLARIS" ``` ### DELETE /api/v1/documentos/{id} Lo borra con todo lo derivado (páginas, vectores, imágenes). ```sh curl -s -X DELETE https://scholaris.joseluissaorin.com/api/v1/documentos/ID -H "Authorization: Bearer $SCHOLARIS" ``` ### GET /api/v1/documentos/{id}/texto Lee el texto por páginas o por un tramo de tiempo. `desde` y `hasta` admiten el folio impreso (`23`, `xiv`), la posición física entre corchetes (`[12]`) y, en audio y vídeo, tiempos (`1:06:56`). ```sh # Las páginas 23 a 25 tal como están impresas curl -sG https://scholaris.joseluissaorin.com/api/v1/documentos/ID/texto -H "Authorization: Bearer $SCHOLARIS" -d desde=23 -d hasta=25 -d formato=markdown # Un tramo de una entrevista curl -sG https://scholaris.joseluissaorin.com/api/v1/documentos/ID/texto -H "Authorization: Bearer $SCHOLARIS" -d desde=1:06:00 -d hasta=1:08:00 ``` ### GET /api/v1/buscar Busca pasajes (búsqueda híbrida y reordenada). Entre comillas, una frase literal. Cada pasaje trae `cita`, `localizador`, `ancla`, `enlace` y el `texto` literal. ```sh curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=la música me metía en el tiempo" -d k=5 ``` ### POST /api/v1/preguntar Responde en Markdown con notas al pie [^n]. El redactor solo puede citar los pasajes que se le dan, y cada nota se comprueba. Con `stream` llega por eventos. ```sh curl -s https://scholaris.joseluissaorin.com/api/v1/preguntar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"pregunta": "¿Qué le pasa a Johnny con el tiempo cuando toca?"}' ``` ### POST /api/v1/citar Devuelve tu texto con las citas insertadas (página exacta y cualquier estilo CSL) y la bibliografía. También acepta un .docx y lo devuelve citado, con su formato. ```sh curl -s https://scholaris.joseluissaorin.com/api/v1/citar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"texto": "La música no saca a Johnny del tiempo: lo mete en otro.", "estilo": "chicago-author-date"}' # Un .docx, devuelto citado curl -s "https://scholaris.joseluissaorin.com/api/v1/citar?estilo=apa" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/vnd.openxmlformats-officedocument.wordprocessingml.document" \ -H "Accept: application/vnd.openxmlformats-officedocument.wordprocessingml.document" \ --data-binary @trabajo.docx -o trabajo-citado.docx ``` ### POST /api/v1/verificar Dice si tu biblioteca respalda una afirmación: `respaldada`, `parcial`, `sin_respaldo` o `contradicha`, con la probabilidad y los pasajes. ```sh curl -s https://scholaris.joseluissaorin.com/api/v1/verificar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"afirmacion": "Johnny Carter dice que la música lo mete en el tiempo."}' ``` Cualquier lectura admite `formato=markdown` (o la cabecera Accept: text/markdown). Lo que tarda (subir y citar) espera por defecto y responde con el resultado. Para no esperar, `esperar=0` o la cabecera `Prefer: respond-async`: la respuesta es un 202 con `progreso_url`. ## Una sesión de verdad Cinco órdenes grabadas tal cual contra la versión local, con los ficheros del banco de pruebas: un cuento de Cortázar en PDF y una clase en audio ya subida. Las respuestas largas están recortadas donde dice […]. ### Subir un PDF y esperar a que esté leído ```sh curl -s "$B/documentos?esperar=120" -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/pdf" --data-binary @cortazar1959perseguidor.pdf | jq '{id, titulo, autores, estado, unidades, referencia}' ``` ```text { "id": "dmuwzohwaoqul34wy", "titulo": "El perseguidor", "autores": [ "Cortázar, Julio" ], "estado": "listo", "unidades": 37, "referencia": "Cortázar, J. (1996). El perseguidor." } ``` ### Buscar, en Markdown ```sh curl -sG "$B/buscar" -H "Authorization: Bearer $SCHOLARIS" --data-urlencode "q=la música me metía en el tiempo" -d k=2 -d formato=markdown ``` ```text # la música me metía en el tiempo ## 1. (Cortázar, 1996, p. 5) *El perseguidor* · p. 5 · [abrir](http://localhost:8795/lector/dmuwzohwaoqul34wy?u=5&f=dmuwzohwaoqul34wy%3At0.8) · `dmuwzohwaoqul34wy:t0.8` > Vaya si lo he oído; vaya si he tratado de escribirlo bien y verídicamente en mi biografía de Johnny. > > —Por eso en casa el tiempo no acababa nunca, sabes. De pelea en pelea, casi sin comer. Y para colmo la religión, ah, eso no te lo puedes imaginar. Cuando el maestro me consiguió un saxo que te hubieras muerto de risa si lo ves, entonces creo que me di cuenta en seguida. La música me sacaba del tiempo, aunque no es más que una manera de decirlo. Si quieres saber lo que realmente siento, yo creo que la música me metía en el tiempo. Pero entonces hay que creer que este tiempo no tiene nada que ver con... bueno, con nosotros, por decirlo así. > […] ``` ### Un pasaje de audio: el localizador es el segundo ```sh curl -sG "$B/buscar" -H "Authorization: Bearer $SCHOLARIS" --data-urlencode "q=vectors as arrows in space" -d k=1 | jq '.pasajes[0] | {cita, localizador, enlace}' ``` ```text { "cita": "(3Blue1Brown, s. f., 0:11)", "localizador": "0:11", "enlace": "http://localhost:8795/lector/dmuwzm8ascxmk4ayv?t=11.3&f=dmuwzm8ascxmk4ayv%3Ac9tp39p" } ``` ### Preguntar ```sh curl -s "$B/preguntar?formato=markdown" -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" -d '{"pregunta":"¿Qué le pasa a Johnny con el tiempo cuando toca?"}' ``` ```text Cuando Johnny toca, experimenta que el tiempo cambia[^1][^2]. Aunque en un principio expresa que «la música me sacaba del tiempo», matiza que en realidad siente que lo «metía en el tiempo», entendiendo que se trata de un tiempo distinto que no tiene que ver con el orden cotidiano[^3]. Niega que se trate de una abstracción; más bien describe que simplemente cambia de lugar, comparando la experiencia con entrar en «un ascensor de tiempo» donde transcurren decenas de pisos en lo que dura una sola frase, mientras asuntos como las hipotecas o la religión dejan de existir para él[^4]. Asimismo, percibe la elasticidad del tiempo mediante la metáfora de una valija en la que, al tocar, es capaz de meter «cientos y cientos de trajes» o «una tienda entera»[^5]. Durante esos momentos musicales, Johnny siente que se abre una puerta y que vuela hacia un estado de seguridad donde queda «sobre todo sin tiempo, sin que después... sin que hubiera después... Por un rato no hubo más que siempre...»[^6]. Sin embargo, esa vivencia dura solo mientras se encuentra sumergido en la música; al dejar de tocar, cae de nuevo «de cabeza» en sí mismo[^6]. Bruno observa que en su ejecución Johnny «siempre está tocando mañana», adelantándose sin esfuerzo al hoy[^3], y viviendo intensamente lo que describe como su «cuarto de hora de minuto y medio», una alteración temporal frente al límite impuesto por los relojes[^1][^7]. [^1]: Cortázar, *El perseguidor* (1996), p. 9. [^2]: Cortázar, *El perseguidor* (1996), p. 5. [^3]: Cortázar, *El perseguidor* (1996), p. 5. [^4]: Cortázar, *El perseguidor* (1996), pp. 5-6. [^5]: Cortázar, *El perseguidor* (1996), p. 6. [^6]: Cortázar, *El perseguidor* (1996), p. 34. [^7]: Cortázar, *El perseguidor* (1996), pp. 16-17. --- Fuentes (confianza alta): […] ``` ### Citar un texto propio ```sh curl -s "$B/citar?formato=markdown" -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" -d '{"texto":"Johnny Carter siente que la música no lo saca del tiempo, sino que lo mete en otro. Ese tiempo no se parece al de los relojes.","estilo":"apa"}' ``` ```text Johnny Carter siente que la música no lo saca del tiempo, sino que lo mete en otro (Cortázar, 1996, p. 5). Ese tiempo no se parece al de los relojes (Cortázar, 1996, pp. 8-9). ## Referencias - Cortázar, J. (1996). El perseguidor. ``` ## Desde JavaScript Con fetch y nada más (Node 20 o el navegador). Diez líneas: subir, buscar y preguntar. ```js import { readFile } from 'node:fs/promises'; const B = 'https://scholaris.joseluissaorin.com/api/v1'; const h = { Authorization: `Bearer ${process.env.SCHOLARIS}` }; const pedir = async (ruta, o = {}) => (await fetch(B + ruta, { ...o, headers: { ...h, ...o.headers } })).json(); const doc = await pedir('/documentos?nombre=articulo.pdf', { method: 'POST', headers: { 'Content-Type': 'application/pdf' }, body: await readFile('articulo.pdf') }); const { pasajes } = await pedir(`/buscar?k=3&q=${encodeURIComponent('atención escalada')}`); for (const p of pasajes) console.log(p.cita, p.texto.slice(0, 80), p.enlace); const r = await pedir('/preguntar', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ pregunta: '¿Qué es la atención multicabeza?' }) }); console.log(r.respuesta); ``` ## Desde Python El SDK trae una fachada de cinco verbos sobre esta misma API. Instálalo con pip (solo necesita requests): ```sh pip install scholaris-sdk ``` ```py from scholaris.api import Scholaris s = Scholaris("sch_…") # o la variable SCHOLARIS_CLAVE doc = s.subir("articulo.pdf") # también una URL for p in s.buscar("atención escalada", k=3): print(p["cita"], p["texto"][:80], p["enlace"]) print(s.preguntar("¿Qué es la atención multicabeza?")["respuesta"]) print(s.citar("La atención sustituye a la recurrencia.")["texto"]) print(s.verificar("El Transformer prescinde de la recurrencia.")["veredicto"]) ``` ## Para agentes Hay dos caminos: el servidor MCP, para los agentes que hablan MCP, y esta API, que cualquier agente con una herramienta HTTP sabe usar. En Claude Code, con una clave que tenga el alcance `mcp`: ```sh claude mcp add --transport http scholaris https://scholaris.joseluissaorin.com/mcp \ --header "Authorization: Bearer sch_…" ``` En Claude (web o escritorio): Ajustes → Conectores → Añadir conector personalizado, con esta dirección. Inicias sesión en Scholaris y concedes acceso; no hace falta clave. ```text https://scholaris.joseluissaorin.com/mcp ``` En Cursor, Windsurf y los demás clientes MCP: ```json { "mcpServers": { "scholaris": { "url": "https://scholaris.joseluissaorin.com/mcp", "headers": { "Authorization": "Bearer sch_…" } } } } ``` Un agente que solo sepa leer la web tiene sus instrucciones en /llms.txt: los verbos, los campos y las reglas para citar sin inventar. ```sh curl -s https://scholaris.joseluissaorin.com/llms.txt ``` 1. Cita solo lo que devuelve la API, y copia `cita` y `localizador` tal cual. 2. Si citas literalmente, copia el `texto` literal; si parafraseas, pon igualmente la cita. 3. Da el `enlace`: es la manera de comprobar la página o el segundo. 4. Si la búsqueda no encuentra nada, dilo; no rellenes el hueco de memoria. ## Cuando algo falla Los errores dicen qué ha pasado en castellano (`mensaje`) y en inglés (`message`), con un código estable para el programa y el enlace a esta guía: | Código | HTTP | Qué significa | | --- | --- | --- | | `no_autenticado` | 401 | Falta la clave, no vale o se ha revocado. | | `prohibido` | 403 | La clave no tiene el alcance necesario (escribir necesita `escritura`). | | `peticion_invalida` | 400 | Falta un campo o no es válido: el mensaje dice cuál. | | `no_encontrado` | 404 | No existe o no es tuyo. | | `conflicto` | 409 | Choca con el estado actual; por ejemplo, el documento aún no tiene texto. | | `demasiado_grande` | 413 | Por aquí caben ficheros de hasta 95 MB; para más, el SDK sube por partes. | | `cuota_superada · requiere_pro` | 402 | Se ha agotado una cuota del plan. | | `limite_de_ritmo` | 429 | Demasiadas peticiones seguidas: espera los segundos de Retry-After. | | `proveedor_fallo` | 502 | Ha fallado un proveedor de IA: vuelve a intentarlo. | Los límites de ritmo y las cuotas son los de tu plan, los mismos que en la aplicación. Para reintentar sin duplicar, manda una cabecera Idempotency-Key en «subir» y «citar»: con la misma clave, durante 24 horas, se devuelve el mismo recurso. ## De dónde sale cada cita Al leer un documento, Scholaris guarda de cada pasaje un ancla: la página física y el folio impreso que se ve en el papel, el segundo de un audio o de un vídeo, la diapositiva, la sección y el párrafo de una web. Es lo que devuelve el campo `ancla`. La `cita` y el `localizador` se escriben desde esa ancla, nunca desde lo que diga un modelo; el `enlace` abre el lector en esa misma página o en ese mismo segundo. Si un pasaje no está en tu biblioteca, la API no lo cita. # ===== ENGLISH ===== --- # Everything worth knowing about Scholaris URL: https://scholaris.joseluissaorin.com/en/knowledge > What Scholaris reads, how it finds and cites, what format it stores, what it does with your data, what it costs and how an agent uses it. More detailed than the front page, and written to be read slowly. ## What this part is for The [front page](https://scholaris.joseluissaorin.com/en.md) says why Scholaris exists. These pages say how it works, with the specifics: which formats go in, where each page number comes from, which provider reads what, how long the things we measured take and what they cost, and what it does not do yet. If anything here disagrees with what you see in the app, the app is right and this page is wrong: write to us and we will fix it. Every page has a Markdown twin (the same address ending in `.md`). There is an index for machines at [/llms.txt](https://scholaris.joseluissaorin.com/llms.txt) and the full text of every page at [/llms-full.txt](https://scholaris.joseluissaorin.com/llms-full.txt). If you are an agent, start with the [page for agents](https://scholaris.joseluissaorin.com/en/agents.md). ## The pages {#hojas} - [What Scholaris is, and what it will never do](https://scholaris.joseluissaorin.com/en/knowledge/what-it-is.md): A library that reads your sources and, when you ask, points to the exact printed page or second. The rules that hold it up: never invent a citation, show the provenance, leave the thinking to whoever writes. - [Which formats go in, and how each citation is anchored](https://scholaris.joseluissaorin.com/en/knowledge/formats.md): Digital and scanned PDFs, photos, EPUB, Word, slides, spreadsheets, audio, video, web pages, YouTube and podcasts. For each, how it is read and which anchor it leaves: printed folio, second, slide, rows or paragraph. - [The SPDF 4.1 format, a library you can take with you](https://scholaris.joseluissaorin.com/en/knowledge/spdf.md): An .spdf is a gzip-compressed SQLite database holding the document already read: text, anchors, sections, figures, vectors and provenance. It opens without Scholaris and offline. - [How it searches, by words, by meaning and by image](https://scholaris.joseluissaorin.com/en/knowledge/search.md): Three searches at once (lexical, semantic and visual), fused and reranked; exact phrases in quotes; a modernised-spelling layer for Old Spanish and Latin; and results that arrive in two stages. - [Citations that are checked before they reach you](https://scholaris.joseluissaorin.com/en/knowledge/citations.md): Nine CSL styles, BibTeX, RIS and CSL-JSON; an autocite that reads your draft and proposes each reference with its page; claim verification with temporal logic; and insertion into .docx without breaking the formatting. - [The tools for thinking with a whole library](https://scholaris.joseluissaorin.com/en/knowledge/research-tools.md): The graph of people, works, places and concepts linked to Wikidata; the concept map; the citation graph between your books; watchers that tell you when something changes; and notebooks whose citations are checked again. - [Libraries that fill up at once and can be shared](https://scholaris.joseluissaorin.com/en/knowledge/libraries.md): Fill a library with folders, ZIPs, lists of links or a Zotero BibTeX; export it as a .scholaris package; share it, follow it, copy it or publish it with a link, respecting its rights. - [Interviews, lectures and videos you can cite to the second](https://scholaris.joseluissaorin.com/en/knowledge/media-player.md): Word-by-word transcription with who is speaking, citable stretches of 30 to 60 seconds, searchable video frames, and a player with a synchronised transcript where you cite by selecting the text. - [API v1 and the MCP server](https://scholaris.joseluissaorin.com/en/knowledge/api-and-mcp.md): Nine verbs over HTTP with one key, answers in JSON or Markdown, and an MCP server with OAuth for Claude, Cursor and other agents. Limits, scopes, errors and examples to copy. - [Scholaris on your computer, or in your own Cloudflare account](https://scholaris.joseluissaorin.com/en/knowledge/self-hosting.md): The same app with SQLite and your disk: on Node, in Docker, as a desktop executable or deployed to your Cloudflare account. What stays on your machine and what still needs the cloud. - [What happens to your data, provider by provider](https://scholaris.joseluissaorin.com/en/knowledge/privacy.md): Where your documents are stored, which AI provider reads what, what is sent to the open bibliographic databases, how your keys are encrypted and what is deleted when you delete. - [Plans, coupons and limits](https://scholaris.joseluissaorin.com/en/knowledge/plans.md): A free plan to start, Pro for real work, and the home version with no quotas. How much fits in each, how pages and minutes are counted, and how coupons work. - [What we measured, and under what conditions](https://scholaris.joseluissaorin.com/en/knowledge/performance.md): Reading times and costs, folio accuracy, search quality, invented citations and latency, as they came out of the benchmark of 6 October 2026, and what those figures do not say. - [Frequently asked questions](https://scholaris.joseluissaorin.com/en/knowledge/faq.md): Whether it invents citations, whether it works with seventeenth-century scans or video interviews, what happens to your data, whether it works offline, and other questions, answered plainly. - [Scholaris glossary](https://scholaris.joseluissaorin.com/en/knowledge/glossary.md): The words Scholaris uses, from folio and anchor to pliego, imprenta, SPDF or watcher, with what they mean here and, where it helps, in the tradition of the book. - [Scholaris beside Zotero, Elicit, NotebookLM and the rest](https://scholaris.joseluissaorin.com/en/knowledge/alternatives.md): What each tool does, when another one suits you better than Scholaris, and what Scholaris does that the others do not set out to do. No figures about others that we cannot check. - [Changelog](https://scholaris.joseluissaorin.com/en/knowledge/changelog.md): What has changed in Scholaris, from the first version to the second, told by milestones and in the order they were made. - [For agents: how to read Scholaris and how to use it on someone’s behalf](https://scholaris.joseluissaorin.com/en/agents.md): What an agent can read on this site and in which format, how to act on a person’s library through API v1 or MCP, with which permissions and limits, and the rules for citing without inventing. ## Scholaris in three sentences - It reads what you give it (digital and scanned PDFs, photos of pages, EPUB, Word, slides, spreadsheets, audio, video, web pages and YouTube) and stores an anchor for every passage: the printed page, the second, the slide or the paragraph. - When you search or ask, it gives you passages with a citation ready to paste, and every citation is written from the stored anchor, never from what a model says. - It does not write for you or decide which author is right: it points to the place and steps aside. ## Who makes it Scholaris is made by [José Luis Saorín Ferrer](https://joseluissaorin.com), philologist and programmer, in Santa Cruz de Tenerife. Write to [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- # What Scholaris is, and what it will never do URL: https://scholaris.joseluissaorin.com/en/knowledge/what-it-is > A library that reads your sources and, when you ask, points to the exact printed page or second. The rules that hold it up: never invent a citation, show the provenance, leave the thinking to whoever writes. ## What it does Scholaris is a web app (and a version for your own computer) for people who read in order to write. You upload your sources: scanned books, papers, theses, notes, slides, recorded interviews, lectures, YouTube videos, web pages. Scholaris reads them once, carefully, and from then on you can: - **search** your whole library at once, by exact words or by meaning, in any language, Old Spanish and Latin included; - **ask**, and get a short answer with footnotes, each tied to its passage; - **cite** in whatever style your journal wants, with the printed page of the edition in front of you or the exact minute of the recording; - **verify** whether your library supports a sentence in your draft; - **explore** what your library holds: people, works, places and concepts, and how the books cite each other. ## The four rules ### Point to the exact page A citation that does not say where it is cannot be checked. So Scholaris stores an **anchor** for every passage: the printed folio you see on paper (not the number your PDF viewer shows), the second of a recording, the slide, the sheet and rows of a table, or the section and paragraph of a web page with its access date. How each anchor is worked out is in [Formats and anchors](https://scholaris.joseluissaorin.com/en/knowledge/formats.md). ### Never invent a citation No citation comes from a language model. A model may draft an answer, but it may only cite the passages it is given, and every note is checked against the text before you see it; the reference (author, year, page) is written from the stored anchor. If there is no passage, the answer says so. In the citation benchmark, over 41 claims, invented citations were zero (see [Performance](https://scholaris.joseluissaorin.com/en/knowledge/performance.md)). ### Show the provenance Every piece of data shows where it came from: which source the record came from (the PDF itself, Crossref, OpenAlex, Open Library, Wikidata), how confidently a page number was read, whether it was printed or inferred, and who corrected it. What is not known is left empty instead of being filled in. ### Leave the thinking to whoever writes Scholaris does not write essays, does not summarise books so you can skip them and does not decide which quotation is the good one. It takes you to the place; what you do with what you find is yours. ## What it does not do (yet, or ever) - It does not search the internet for you: it works on what you give it. For discovering new literature there are better tools (see [Alternatives](https://scholaris.joseluissaorin.com/en/knowledge/alternatives.md)). - It does not write your text. The "cite" function gives you back your own text with the citations inserted, not a new text. - It cannot guarantee a printed page number when the book does not print one: in that case it cites the physical position in brackets, "p. [12]", and says so. - It does not work fully offline: even the home version needs a cloud provider to read pages (see [Self-hosting](https://scholaris.joseluissaorin.com/en/knowledge/self-hosting.md)). ## Who it is for For the doctoral student with three hundred PDFs and a supervisor who wants the page number; for the lecturer preparing a class with twenty years of underlining; for the journalist going back to a two-year-old interview in search of one sentence; for translators, archivists and students, and for anyone who has ever said "I read it somewhere". --- # Which formats go in, and how each citation is anchored URL: https://scholaris.joseluissaorin.com/en/knowledge/formats > Digital and scanned PDFs, photos, EPUB, Word, slides, spreadsheets, audio, video, web pages, YouTube and podcasts. For each, how it is read and which anchor it leaves: printed folio, second, slide, rows or paragraph. ## How a file goes in Everything first goes through the **imprenta** (the press), Scholaris's converter. It recognises the type from the first bytes of the file (then from its MIME type and, last, its extension) and converts it in your own browser whenever it can. What the browser cannot do (decode some images, extract video frames, convert a Keynote deck) is done by a server container with ffmpeg and LibreOffice. Then whatever has to be read with eyes (scanned pages, photos, slides) is read by a vision model, four pages per request, and whatever has to be heard is transcribed by a speech model. Which provider does what is in [Privacy](https://scholaris.joseluissaorin.com/en/knowledge/privacy.md). ## Formats and anchors | What you upload | Extensions | How it is read | Anchor and how it is cited | | --- | --- | --- | --- | | Digital PDF | pdf | The text layer, with its lines and blocks; headers and footers set apart; the PDF's own page labels, outline and metadata are used | Printed page: "p. 23", "pp. 23-24", "p. xiv" | | Scanned PDF, or one with old OCR | pdf | Vision: each page as an image. If the PDF carries a poor old OCR layer, it is read again | Printed page, or the physical one in brackets: "p. [12]" | | Photos of a book | jpg, png, webp, gif, tiff, heic, bmp, avif | Sorted by name in numeric order (IMG_2 before IMG_10), straightened and cropped | Printed page of each photo | | Text documents | docx, odt, rtf, html, md, txt | Blocks (headings, paragraphs, lists, quotations, tables, notes) with their heading path | Section and paragraph: "Introduction, para. 4" | | EPUB | epub | Reading order, table of contents and, above all, the print page list when the EPUB has one | "p. 23" with a page list; otherwise "para. N" | | Slides | pptx, odp, key | Text and speaker notes of every slide; ODP and Keynote are converted whole with LibreOffice | "slide 7" | | Spreadsheets | xlsx, xls, ods, csv, tsv | Tables in chunks of 50 rows, with the header repeated in each chunk | "Sheet1, rows 2-51" | | Audio | mp3, wav, m4a, ogg, opus, flac, aac | Word-by-word transcription with who is speaking; citable stretches of 30 to 60 seconds that end at a sentence boundary | "12:04", "1:02:03", "12:04-12:40" | | Video | mp4, mov, webm, mkv, avi | Audio as above; plus frames at every scene change (or every 20 seconds), which are searchable too | Same as audio | | A web page | URL | A dated copy is stored and the article is split into blocks | Section and paragraph, with the access date | | YouTube | URL | Not downloaded: Gemini watches the video from its address, ten minutes at a time | The second | | Vimeo, podcasts, audio links | URL | Vimeo, the smallest progressive file; a podcast, the audio in its RSS | The second | | SPDF and .scholaris packages | spdf, scholaris | Nothing is read again: text, anchors and vectors are copied | The anchors they carried | ## The printed folio The number that matters in a footnote is the one printed on paper, not the page's position in the file. A book may start in roman numerals, skip unnumbered plates or have two pages per image. For each page Scholaris gathers several pieces of evidence: 1. the PDF's own page labels, if any; 2. candidate numbers from the header, the footer and the text edges, plus what the vision reader saw; 3. the longest coherent sequence of those readings; 4. for doubtful cases, a judge (Jev, by TypeSafe) choosing between candidates; 5. and, for pages with no visible number, interpolation within each stretch. It recognises roman numerals (and the page where arabic numbering starts), foliation by leaves ("23r", "23v"), two-up scans, unnumbered plates and years that are not folios. Every page records where its number came from (read, inferred, from the EPUB, or none) and how confident it is. **If a book prints no number at all, none is invented:** the physical position is cited in brackets. In the benchmark, all 64 folios checked by eye came out exact (see [Performance](https://scholaris.joseluissaorin.com/en/knowledge/performance.md)). ## Original spelling is kept The reader transcribes faithfully: it does not modernise "dixo", "assi" or "muger", keeps u/v, i/j/y and ç, does not fix misprints and keeps running heads, catchwords and signatures out of the body. The one exception is the long s (ſ), written as "s". So that those spellings can still be found, there is a separate layer, explained in [Search](https://scholaris.joseluissaorin.com/en/knowledge/search.md). ## Size limits | Plan | Largest file | Storage | | --- | --- | --- | | Free | 200 MB | 1 GB | | Pro | 4 GB | 100 GB | | Home version | 16 GB | your disk | Through API v1 in the cloud a single request takes files up to 95 MB; for more, the app and the Python SDK upload in parts. There is no maximum duration for audio and video beyond the file size. ## What does not work well yet - Old Word .doc files have no reader of their own: convert them to .docx first. - Vimeo only works when the video offers a downloadable file; otherwise, download it and upload it. - Web anchors are paragraph anchors: if the page changes later, the dated copy we stored is the one that counts. --- # The SPDF 4.1 format, a library you can take with you URL: https://scholaris.joseluissaorin.com/en/knowledge/spdf > An .spdf is a gzip-compressed SQLite database holding the document already read: text, anchors, sections, figures, vectors and provenance. It opens without Scholaris and offline. ## What an .spdf is An **.spdf** file (Scholaris PDF) stores a document already read: not just the original but everything Scholaris got out of it, so it can be searched and cited elsewhere without reading it again or paying for it twice. Inside it is a **gzip-compressed SQLite database** (not a ZIP). Any language with SQLite opens it; uncompressed files are accepted too. The current version is **4.1**; the version is stored in the `spdf` table under the key `spdf_version`. Old SPDFs (versions 1 to 3, from the first Scholaris) are migrated on the fly when opened. Each exported .spdf holds one document. The same schema is what each user's library uses on the server. ## The tables | Table | What it stores | | --- | --- | | spdf | Key and value: version, creation date, generator, hash of the original | | documentos | Type, record (JSON), SHA-256 hash, MIME type, size, number of units, duration, title, authors, year and language | | unidades | The citable units (pages, audio stretches, slides): anchor, text, notes, header and footer, image, confidence, printed page, times and per-word timings | | secciones | The section tree | | fragmentos | The passages that are searched and cited: text, context, section, start and end anchors, and the modernised-spelling layer (texto_busqueda) | | fragmentos_fts | FTS5 full-text index over text, context, section and modernised spelling, accent-insensitive | | figuras | Figures, plates and video frames, with region, caption and description | | espacios | Vector spaces: provider, model, version, dimensions, normalisation and modalities | | vectores | One vector per target (fragment, unit or figure) and space, as little-endian float32 | | blobs | The original and the images (only in exported files) | | procedencia | The log of how it was read: each phase, which provider did it, how long it took and when | Version 4.1 added the `texto_busqueda` column: a shadow of the text in modernised spelling, used only for searching (see [Search](https://scholaris.joseluissaorin.com/en/knowledge/search.md)). The text that gets cited is never touched. ## Vector spaces One .spdf can carry vectors from several models at once, each declared in the `espacios` table. The ones Scholaris uses today: | Space | Model | Dimensions | When | | --- | --- | --- | --- | | gemini-embedding-2@1536 | Gemini Embedding 2, multimodal, truncated | 1536 | By default, in the cloud and at home | | qwen3-vl-embedding-2b@2048 | Qwen3-VL Embedding 2B, on InferBox | 2048 | In the home version with a GPU | | qwen3-embedding-0.6b@1024 | Qwen3 Embedding 0.6B, on Workers AI | 1024 | When there is no Gemini key | When an .spdf is imported, only the vectors missing for your library's space are computed; text and anchors are not read again. ## Opening it offline with Python The Python SDK includes an SPDF reader that uses only the standard library (gzip and sqlite3) and opens the file read-only: ```sh pip install scholaris-sdk # only needs requests pip install "scholaris-sdk[vectores]" # with numpy, for vector search ``` ```py from scholaris.v2 import SPDF with SPDF.abrir("vigilar.spdf") as s: print(s.documento["metadatos"]["titulo"]) for f in s.buscar("panóptico", k=5): # FTS5, accent-insensitive print(f.cita, f.texto[:80]) # «(Foucault, 1975, p. 23) …» print(s.pagina("145").texto) # by printed page number ``` The reader also exposes `documentos`, `unidades`, `fragmentos`, `secciones`, `espacios`, `blob`, `original` and `vectores(espacio)`, and does similarity search with `buscar_vector(vector, espacio, k)`. The package is `scholaris-sdk` on PyPI and is imported as `scholaris`; it is licensed under the EUPL-1.2. One caveat: the Python reader searches the FTS5 index as is, without the modernised-spelling layer that Scholaris applies to the query. ## Opening it without Python ```sh gzip -dc libro.spdf > libro.sqlite sqlite3 libro.sqlite "SELECT valor FROM spdf WHERE clave = 'spdf_version'" sqlite3 libro.sqlite "SELECT texto FROM fragmentos_fts WHERE fragmentos_fts MATCH 'panoptico' LIMIT 3" ``` ## Export and import - From the app, every document exports as an .spdf, with or without the original and the vectors. The server embeds up to 24 MB of files; for more, the export runs in the browser. - A whole library exports as a **.scholaris** package: a ZIP with one .spdf per document, a `manifest.json` (format `scholaris-biblioteca`, version 1) and a `LEEME.txt` with the library's rights. See [Libraries](https://scholaris.joseluissaorin.com/en/knowledge/libraries.md). - Importing an .spdf (up to 512 MB) reads nothing again. ## What is missing There is no public validator for version 4 yet. The full specification lives in the code (`packages/spdf/esquema/v4.1.sql`), which will be published with the rest of the source (see [Self-hosting](https://scholaris.joseluissaorin.com/en/knowledge/self-hosting.md)). --- # How it searches, by words, by meaning and by image URL: https://scholaris.joseluissaorin.com/en/knowledge/search > Three searches at once (lexical, semantic and visual), fused and reranked; exact phrases in quotes; a modernised-spelling layer for Old Spanish and Latin; and results that arrive in two stages. ## Three routes that meet Every query goes down three routes at once, and the results are fused by reciprocal rank (RRF), with weights that depend on what you seem to be looking for: | Route | What it finds | How | | --- | --- | --- | | Lexical | The words you type, accent-insensitive | SQLite FTS5 index with BM25; section headings and context weigh less than the text | | Semantic | What the query means, even when the passage uses other words or another language | Gemini Embedding 2 vectors (1536 dimensions) in Cloudflare Vectorize | | Visual | Plates, figures, charts and video frames | The same multimodal vectors, applied to page images and figures | | If the query looks like… | lexical | semantic | visual | | --- | --- | --- | --- | | an idea | 0.35 | 1 | 0 | | a literal quotation | 1 | 0.45 | 0 | | something you can see | 0.35 | 0.6 | 1 | | a question about dates | 0.35 | 1 | 0 | The top thirty candidates are reranked by a judge, Jev (by TypeSafe), which answers for each passage whether it addresses the query; the final score is 20 % fusion and 80 % judge. The visual route only counts when you are looking for something visual: on its own it does poorly on text questions and was making them worse. ## Exact phrases In quotes ("…", “…” or «…») the phrase is searched literally, with no model involved. Exact matches come first, then passages close in meaning. If the phrase appears nowhere, Scholaris says so ("La frase exacta no aparece; se busca por sentido": the exact phrase does not appear, searching by meaning) instead of pretending it found it. ## Old Spanish and Latin {#grafia} A seventeenth-century book writes "dixo", "muger", "assi" or "agora"; you search for "dijo", "mujer", "así", "ahora". So that they meet without touching the text that is cited, every passage carries a **modernised-spelling shadow** used only for searching, and the query is reduced with the same rules: - **Spanish, 16th to 18th century:** word lists (fee → fe, agora → ahora, mesmo → mismo, truxo → trajo) and phonetic and spelling rules (b/v, c/z, g/j, h, x → j, ph, rr…); - **Latin:** u/v, i/j, æ/œ, michi → mihi, the enclitics -que/-ne/-ve and the Schinke et al. (1996) stemmer; - **Old French and Italian:** basic rules. The layer always applies to Latin, and to the other languages when the document predates a cut-off (1830 for Spanish, 1800 for French and Italian) or the text itself gives signs (ſ, ç, "assi", "dixo"…). It is rules, no models. On a play by Lope de Vega, average recall per query went from 32 % to 100 %; on modern documents, searches return exactly the same as without the layer (see [Performance](https://scholaris.joseluissaorin.com/en/knowledge/performance.md)). ## Results in two stages In the app, search arrives in three deliveries over an event stream: first what the lexical index finds, without waiting for the query vector; then the fusion of the three routes, at around 350 ms; and finally the judge's order, at around 650 ms median measured from home. What you see first is already useful and does not get reshuffled all at once. ## Filters You can filter by library, document, type, author, language and year (from, to). Some filters are understood when written into the query itself: "before 1980", "since 1950", the surname of an author in your library. "Go to page 145" jumps to that printed page. ## What each result returns The literal passage, its citation ready to paste ("(Foucault, 1975, p. 23)"), the exact locator, the anchor and a link that opens the reader at that page or second. Search never writes a citation: the citation comes from the anchor. ## Also - **Ask:** a short Markdown answer with footnotes; the writer may only cite the passages it is given, and every note is checked. - **Similar:** passages and figures similar to a given one. - **Across languages:** the query translated into the languages of your library. - **Libraries you follow:** search your own and the ones others share with you at the same time. --- # Citations that are checked before they reach you URL: https://scholaris.joseluissaorin.com/en/knowledge/citations > Nine CSL styles, BibTeX, RIS and CSL-JSON; an autocite that reads your draft and proposes each reference with its page; claim verification with temporal logic; and insertion into .docx without breaking the formatting. ## Upload and cite, step by step {#subir-y-citar} 1. Sign in to Scholaris, or try it without an account in [the demo](https://scholaris.joseluissaorin.com/?demostracion). 2. In the Library (Biblioteca), drop the file (or paste an address: a web page, a PDF, YouTube, a podcast). The first pages of a digital PDF are searchable in under a second; those of a scan, in a few seconds. 3. In Search (Buscar), type the idea in your own words or the exact phrase in quotes. 4. In the result, copy the reference: it comes out in your style, with the printed page or the minute, as formatted text (italics included). 5. For a whole text, go to Write (Escribir) and paste your draft or upload your .docx: autocite proposes a citation for each claim, with its page, and you accept or discard them one by one. 6. Export the cited text (.docx, Markdown, plain text or LaTeX) and the bibliography (BibTeX, RIS or CSL-JSON). ## Styles Citations are formatted by citeproc-js with the official CSL files, in four locales (es-ES, en-US, fr-FR, it-IT). The default style is APA. | Identifier | Style | | --- | --- | | apa | APA, 7th edition | | chicago-author-date | Chicago, 18th edition, author-date | | chicago-note-bibliography | Chicago, 18th edition, notes and bibliography | | mla | MLA, 9th edition | | harvard | Harvard (Cite Them Right, 12th edition) | | iso690 | ISO 690 author-date, Spanish | | iso690-en | ISO 690 author-date, English | | iso690-numerico | ISO 690 numeric | | ieee | IEEE | When the date of the original work differs from the edition's, it appears as the original date ("1605/2004"); the edition is given unless it is a same-year reprint. ## Where each citation comes from No model writes the citation. The locator ("p. 23", "12:04", "slide 7") comes from the anchor stored when the document was read (see [Formats and anchors](https://scholaris.joseluissaorin.com/en/knowledge/formats.md)), and author, year and title come from the record, whose provenance you can inspect field by field. The record is checked against Crossref, OpenAlex, Open Library and Wikidata, and an external record is only accepted if the title matches and the author or the year matches too. ## Autocite For a whole draft, Scholaris: 1. splits it into paragraphs and claims; 2. retrieves fourteen candidate passages for each; 3. lets a model propose which passage supports each claim; 4. checks in code that the evidence appears verbatim in the passage, that a negative statement is not cited as if it were positive, that the key terms are there and that the chronology is possible; 5. and runs every proposal past a judge (Jev, by TypeSafe) in batches of eight. Proposals with probability 0.7 or more are accepted; between 0.4 and 0.7, flagged for review; below that, discarded. Consecutive verified pages are merged ("pp. 23-25"). In the benchmark, precision was 95.1 %, recall 96.7 % and invented citations zero. ## Verifying a claim You give it a sentence and Scholaris says whether your library supports it, with the passages: respaldada (supported) : the best passage supports it directly, with probability 0.7 or more. parcial (partial) : some passage partly supports it (probability 0.4 or more). contradicha (contradicted) : some passage says the opposite, with probability 0.5 or more. sin_respaldo (unsupported) : no passage supports it. The judge also classifies how each passage relates to the claim: direct support, application of a framework, context, contradiction, reported opinion or negative statement. ## Temporal logic Comparing dates is not left to a model: code does it. - The year of the original work is used, not the edition's. - A text cannot directly support a claim about something invented later: Scholaris knows when terms like "internet" (1983), "deconstruction" (1967) or "transformer" (2017) appeared. - If the claim says one author "anticipated" or "built on" another, the chronology has to be possible; if it is not, the citation is dropped. - If the claim refers to a year more than twenty years after the source, direct support is downgraded to "application of a framework". - A source later than the year of your own text cannot be cited in it. - Undated sources ("n.d.") skip all of this. ## Inside Word Scholaris opens the .docx, inserts each citation in place keeping the formatting of every run of text, uses real Word footnotes for note styles and appends the bibliography. It also takes ODT, plain text, Markdown and HTML, up to 50 MB. ## Copying a reference Every document has a button to copy its reference in your style (or another, with a preview). It is copied as formatted text, italics included, and as plain text. ## Export and import - Export: BibTeX, RIS and CSL-JSON; the bibliography as text, HTML or Markdown. - Import a BibTeX file (from Zotero, say): each entry is matched to your documents by DOI, then ISBN, then title; those that do not match stay as references without text. --- # The tools for thinking with a whole library URL: https://scholaris.joseluissaorin.com/en/knowledge/research-tools > The graph of people, works, places and concepts linked to Wikidata; the concept map; the citation graph between your books; watchers that tell you when something changes; and notebooks whose citations are checked again. ## People, works, places and concepts Scholaris reads the whole library looking for entities: people, works, places, organisations, concepts, events and dates (and it tells fictional characters apart). A model proposes them in batches of about 48,000 characters, and then code finds each mention in the text, so every occurrence has its exact anchor and can be opened at its page. Each entity is linked, where possible, to its Wikidata record. Two entities are joined when they appear close together: the edge weight falls with distance inside a passage and is lower between neighbouring passages. For the strongest relations in each document, a model names the relation. From the graph you can see an entity's neighbours, its timeline and the shortest path between two. Measured in the benchmark, extracting the entities of 300 pages costs about $0.03. ## The concept map A view of the library as a chart of the sky: the vectors of passages and figures are clustered, projected to two dimensions with UMAP (on a sample of up to 8,000 points) and each cluster gets a name. It shows what your library talks about and what lies far from what. Separately there is **Concepts**: you type a concept and Scholaris gathers where it is defined, applied, criticised and merely mentioned, each with its passage and page. It exports to CSV, JSON, JSONL, HTML, BibTeX, TEI and XLSX. ## The citation graph How the documents in your library cite each other. The bibliography entries of each document are resolved against your library by DOI, by title (with a compatible year) or by first author and year when there is a single candidate, and every link carries its confidence. What your books cite and you do not have shows up as **orphan references**: a list of what you still have to read. The graph only looks inside your library; it does not query external databases. ## Watchers A watcher is a saved search that repeats itself: by hand, on every upload, daily or weekly (in the cloud, every day at 06:17 UTC, and the weekly ones on Mondays). It tells you when a new document enters the top results or when the answer changes: if confidence rises or falls, or if it says something else. ## Notebooks Markdown notes with cards: a passage, a page, a figure, a stretch of audio, a note of yours or a synthesis. Each card keeps a verified copy of the passage, its anchor and a hash, and with **verify again** Scholaris checks that the library still says the same thing: living citations. A notebook's synthesis may only cite the notebook's own cards; any other mark is removed. ## Perspectives Documents you have not opened in a while, gaps (topics you have little on) and recommendations from within what you already have. --- # Libraries that fill up at once and can be shared URL: https://scholaris.joseluissaorin.com/en/knowledge/libraries > Fill a library with folders, ZIPs, lists of links or a Zotero BibTeX; export it as a .scholaris package; share it, follow it, copy it or publish it with a link, respecting its rights. ## Filling a library in one go From "Llenar biblioteca" (fill library) you can drop a whole folder, many files, a ZIP, a list of links, a BibTeX or RIS file with its Zotero folder beside it, or a pile of .spdf files. Before starting, Scholaris shows how long it will take and what it will cost, and lets you choose between two modes: | Mode | What it does | For | | --- | --- | --- | | Fast | Reads right away | When you need it today | | Economy | Hard pages go through Gemini's batch API at half price and easy ones through a cheaper reader; takes hours | Filling a large library | Repeated files (same SHA-256 hash) are detected and not read twice. A batch can be paused and resumed. ## .scholaris packages A whole library exports as a **.scholaris** file: a ZIP with one [.spdf](https://scholaris.joseluissaorin.com/en/knowledge/spdf.md) per document, a `manifest.json` with the library, its rights and its documents, and a `LEEME.txt` (read-me). Originals and vectors are optional. Importing it into another account, or into the home version, reads nothing again. ## Sharing - **Invite** by email, with read, edit or admin permission, and an expiry (from one day to ten years). - **Follow** a library shared with you, and search it alongside your own. - **Copy** into your library: the copy is by reference, so no files are duplicated and nothing is read again. - **Public read-only link**, with an optional password, expiry, visit counter and the option to revoke it. Those pages ask search engines not to index them. ## Rights Every library declares its rights: unspecified, public domain, CC0, CC BY, CC BY-SA, CC BY-NC, CC BY-NC-SA, private use or with permission. To publish a library that is not open through a link you must confirm it is for private use, and the package of a non-open library carries a do-not-redistribute notice. Sharing a copyrighted book with someone who should not have it is the sharer's responsibility; Scholaris only asks you first. --- # Interviews, lectures and videos you can cite to the second URL: https://scholaris.joseluissaorin.com/en/knowledge/media-player > Word-by-word transcription with who is speaking, citable stretches of 30 to 60 seconds, searchable video frames, and a player with a synchronised transcript where you cite by selecting the text. ## The transcript Audio is transcribed in ten-minute stretches with a two-second overlap. When there is a Gemini key, Gemini Transcribe does it, telling speakers apart; if it fails, Whisper (large-v3-turbo, on Workers AI). Then a model identifies the cast (who speaks and in what role: interviewer, guest, host) and assigns every sentence to a named person, which fixes labels that do not match from one stretch to the next. In two interviews from the benchmark, attribution was right in 20 of 21 turns and in 23 of 23. The text is cut into **citable stretches** of 30 to 60 seconds (45 on average) that end at a sentence boundary. Each stretch is a unit with its time anchor and speaker, and per-word timings are stored too. ## Frames In videos a frame is taken at every scene change and, if nothing changes, one every 20 seconds (never more than one every 10). Frames are described and vectorised, so you can search for "the blackboard with the diagram" and land on the exact second. ## YouTube, Vimeo and podcasts A YouTube video is not downloaded: Gemini watches it from its address, ten minutes at a time, and the title and channel come from its public record. From Vimeo the smallest downloadable file is taken; from a podcast, the audio its RSS announces. ## The player - A timeline with each speaker's turns, the chapters and a frame preview on hover. - A synchronised transcript that highlights the word being spoken; click a word to jump there. - To cite, select the text: the citation comes with its exact interval ("12:04-12:40"). - A small player keeps playing while you move around the rest of the app. ## What we measured A 54-minute video interview was ready in 45 seconds ($0.41), and a two-hour one in 70 seconds ($0.94); transcription costs about $0.005 a minute. In the first Scholaris, the same 54-minute interview took three hours and 39 minutes. All figures and their conditions are in [Performance](https://scholaris.joseluissaorin.com/en/knowledge/performance.md). ## What is missing Per-word timings are estimated within each segment, not measured one by one. Some videos with uncommon codecs are not yet converted to a format every browser can play. --- # API v1 and the MCP server URL: https://scholaris.joseluissaorin.com/en/knowledge/api-and-mcp > Nine verbs over HTTP with one key, answers in JSON or Markdown, and an MCP server with OAuth for Claude, Cursor and other agents. Limits, scopes, errors and examples to copy. ## In one sentence Everything the app does with your documents can be done from outside with an `Authorization: Bearer sch_…` header and nine verbs. The full guide, with a real recorded session, is at [/en/api](https://scholaris.joseluissaorin.com/en/api.md); the specification, in [OpenAPI 3.1](https://scholaris.joseluissaorin.com/api/v1/openapi.json); and the instructions for models, at [/api/v1/llms.txt](https://scholaris.joseluissaorin.com/api/v1/llms.txt). Field names are in Spanish. ## The verbs Base: `https://scholaris.joseluissaorin.com/api/v1` | Method | Path | What it does | | --- | --- | --- | | POST | /documentos | Upload a file (raw body, multipart field «archivo» or JSON with «url») | | GET | /documentos | List the library (q, estado, cursor; up to 200 per page) | | GET | /documentos/{id} | The record; with esperar=30 it waits until it is read | | DELETE | /documentos/{id} | Delete it with everything derived | | GET | /documentos/{id}/texto | Text by printed page, [physical page] or time range | | GET, POST | /buscar | Search passages (k up to 50) | | POST | /preguntar | Answer with checked footnotes; stream by events | | POST | /citar | Your text (or .docx) back with citations and bibliography | | POST | /verificar | Does your library support a claim? | Every read accepts `formato=markdown` or the header `Accept: text/markdown`. What takes time (uploading and citing) waits by default; with `esperar=0` or `Prefer: respond-async` it answers 202 with `progreso_url`. ```sh export SCHOLARIS=sch_… # Settings → API keys # Upload (waits until read, up to 60 s; longer, 202 with progreso_url) curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos?nombre=articulo.pdf" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/pdf" --data-binary @articulo.pdf # Search, in Markdown curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=atención escalada" -d k=5 -d formato=markdown # Verify a claim curl -s https://scholaris.joseluissaorin.com/api/v1/verificar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"afirmacion": "El Transformer prescinde de la recurrencia."}' ``` ```py from scholaris.api import Scholaris s = Scholaris("sch_…") # or the SCHOLARIS_CLAVE variable doc = s.subir("articulo.pdf") # a URL works too for p in s.buscar("atención escalada", k=3): print(p["cita"], p["texto"][:80], p["enlace"]) ``` ## Keys and scopes Keys are created in Ajustes (Settings) → Claves de API (API keys) and shown only once; Scholaris stores only their SHA-256 hash. They can expire after as many days as you choose. Each key has scopes: lectura (read) : search, read, ask and verify. escritura (write) : upload and delete documents, and cite a whole text (autocite stores its work). mcp : use the key with the MCP server. Without explicit scopes, a new key gets lectura and mcp. ## Limits | | Free | Pro | | --- | --- | --- | | Requests per minute | 120 | 600 | | Searches per day | 100 | 5,000 | | Autocites per month | 5 | 500 | | Pages or minutes read per month | 1,500 | 60,000 | | File per API v1 request | 95 MB | 95 MB | Going over the rate gives a 429 with `Retry-After`; using up a quota, a 402 (`cuota_superada` or `requiere_pro`). To retry without duplicating, send `Idempotency-Key` when uploading and citing: with the same key, for 24 hours, you get the same resource back. An identical file (same hash) returns the existing one, with `duplicado: true`. ## Errors They all have the same shape, with the message in Spanish and English: ```json { "error": { "codigo": "prohibido", "mensaje": "Esta clave de API es de solo lectura.", "message": "This key is not allowed to do that.", "estado": 403, "documentacion": "https://scholaris.joseluissaorin.com/api#errores" } } ``` Codes: `no_autenticado` (401), `prohibido` (403), `peticion_invalida` (400), `no_encontrado` (404), `conflicto` (409), `demasiado_grande` (413), `cuota_superada` and `requiere_pro` (402), `limite_de_ritmo` (429), `proveedor_fallo` (502), `no_disponible` and `interno`. ## The MCP server At `https://scholaris.joseluissaorin.com/mcp`, over HTTP (stateless Streamable HTTP). It offers four tools: | Tool | What it does | | --- | --- | | search | Hybrid search; passages with their exact locator | | cite | A CSL citation for a fragment or a document, in the requested style and language | | open_page | The full text of a page, by physical position or printed folio | | verify_claim | A verdict on a claim, with the passages that support or contradict it | There are two ways to connect: - **OAuth 2.1**, for Claude (web and desktop) and any client that speaks it: Settings → Connectors → Add custom connector, with the address `https://scholaris.joseluissaorin.com/mcp`. The client registers itself, you sign in to Scholaris and grant access. That access is read-only: search, open pages, cite and verify; it cannot upload, change or delete. Tokens last an hour and are refreshed. - **A key** with the `mcp` scope, for Claude Code, Cursor, Windsurf and the rest: ```sh claude mcp add --transport http scholaris https://scholaris.joseluissaorin.com/mcp \ --header "Authorization: Bearer sch_…" ``` ```json { "mcpServers": { "scholaris": { "url": "https://scholaris.joseluissaorin.com/mcp", "headers": { "Authorization": "Bearer sch_…" } } } } ``` The OAuth metadata is where clients look for it: [/.well-known/oauth-authorization-server](https://scholaris.joseluissaorin.com/.well-known/oauth-authorization-server) and [/.well-known/oauth-protected-resource/mcp](https://scholaris.joseluissaorin.com/.well-known/oauth-protected-resource/mcp). There is also a server card at [/.well-known/mcp/server-card.json](https://scholaris.joseluissaorin.com/.well-known/mcp/server-card.json). ## For agents The rules of use (cite only what the API returns, copy the citation verbatim, give the link, say when nothing is found) are on the [page for agents](https://scholaris.joseluissaorin.com/en/agents.md). --- # Scholaris on your computer, or in your own Cloudflare account URL: https://scholaris.joseluissaorin.com/en/knowledge/self-hosting > The same app with SQLite and your disk: on Node, in Docker, as a desktop executable or deployed to your Cloudflare account. What stays on your machine and what still needs the cloud. ## Four ways The instance at [scholaris.joseluissaorin.com](https://scholaris.joseluissaorin.com/en.md) is the hosted version, the one paid for with the Pro plan. The same code also runs on your computer: | How | What it is | Where your data lives | | --- | --- | --- | | Node | The same API on Node, SQLite (with sqlite-vec for vectors) and the disk, with a queue that resumes if interrupted; listens on port 8790 | The folder you choose | | Docker | The same in a container with ffmpeg; a variant adds InferBox, a model server with an NVIDIA GPU | A Docker volume | | Desktop | A single executable (Bun) for macOS, Windows and Linux, with the web app inside, that opens the browser on start | ~/Scholaris | | Your Cloudflare account | A script creates the database, storage, vector index and queue, and deploys the Worker (needs paid Workers) | Your account | The home version has no quotas: the "local" plan does not limit documents, pages or searches, and takes files up to 16 GB. It can have a single user, several without an external account, or use Clerk to sign in. ## Open source The code will be published under the **EUPL-1.2** at [github.com/joseluissaorin/scholaris-v2](https://github.com/joseluissaorin/scholaris-v2). While the review is finished the repository is private: we are not giving a date. The exact installation commands will be in its README, which takes precedence over this page. The Python SDK already carries the same licence. ## What stays on your machine and what does not In the home version, **your files, your library, the vectors and the index live on your disk**. But Scholaris does not work fully offline: to read pages it needs at least one cloud reader. | Piece | Offline | With a cloud provider | | --- | --- | --- | | Reading pages (scans, photos, slides) | No local reader yet | Gemini, OpenRouter (Mistral OCR) or Workers AI | | Vectors | InferBox (Qwen3-VL Embedding, 2048 dimensions) | Gemini Embedding 2 | | Reranking and judging | InferBox | Jev (TypeSafe) or Workers AI | | Transcribing | InferBox | Gemini Transcribe or Whisper on Workers AI | | Drafting answers | InferBox | Gemini or OpenRouter | | Searching and citing what is already read | Yes: word search always; search by meaning, with InferBox | | In practice the Gemini key is the only required one; OpenRouter, TypeSafe, Workers AI and OpenAlex are optional. What does work fully offline is **opening and searching an .spdf that has already been read** with the Python SDK (see [The SPDF format](https://scholaris.joseluissaorin.com/en/knowledge/spdf.md)). ## The same from outside The home version speaks the same API (v1 and v2) and the same MCP as the cloud, so the Python SDK, the examples in the [API guide](https://scholaris.joseluissaorin.com/en/api.md) and agents work the same pointed at `http://localhost:8790`. --- # What happens to your data, provider by provider URL: https://scholaris.joseluissaorin.com/en/knowledge/privacy > Where your documents are stored, which AI provider reads what, what is sent to the open bibliographic databases, how your keys are encrypted and what is deleted when you delete. ## The main thing Your documents are yours. We do not train models on them. To read them, transcribe them and search them, Scholaris sends pages, audio and queries to AI providers, and this page says which ones, what for and what is stored where. If you work with material that must not leave your computer, the [home version](https://scholaris.joseluissaorin.com/en/knowledge/self-hosting.md) keeps everything on your disk, although it still needs a provider to read pages. There are no third-party analytics and no ads on the site. The only cookies are the session ones (Clerk) and one that remembers you are in the demo. ## Where each thing is stored Everything lives on Cloudflare, in Scholaris's account: | What | Where | | --- | --- | | Your account, your API keys (hash only), settings, usage, invitations and access log | D1, Cloudflare's database | | Originals and page images | R2, Cloudflare's object storage, under a folder of your own | | Your read library (text, anchors, records, index) | A Durable Object of your own, with its SQLite database, just for you | | Vectors | Cloudflare Vectorize, in a namespace of your own | | The session | Clerk, which handles sign-in (email and name) and payments | ## What each provider reads | Provider | What it receives | What for | | --- | --- | --- | | Google Gemini (paid API) | Page images, audio, the address of YouTube videos, passages and queries | Reading pages (Gemini 3.8 Flash and 3.5 Flash-Lite), transcribing (Gemini Transcribe), vectors (Gemini Embedding 2), drafting answers and extracting entities | | Cloudflare Workers AI | Audio, images of easy pages and text | Transcribing with Whisper, fallback and economy reader, fallback vectors and reranker | | OpenRouter (with Mistral OCR) | Pages the previous readers could not read; text to draft if Gemini fails | Fallback reader and writer | | TypeSafe (Jev) | Query and passage pairs, claims and passages, page-number candidates | Reranking results, judging citations and settling doubtful folios | These providers get just what each task needs and are used through their paid or business APIs. Their terms (not ours) say how long they keep data and whether they use it for anything else; for instance, the Gemini API's paid terms exclude using requests to improve their products, while allowing them to be kept for a limited time to detect abuse. If that is not enough for you, use the home version with your own keys or with InferBox. ## What is sent to open bibliographic databases To complete and check each document's record, Scholaris asks Crossref, OpenAlex, Open Library, Wikidata, Wikipedia, arXiv, DataCite and, as a last resort, Google Books. What is sent is **the title, the authors, the ISBN or the DOI**, never the text. To link entities, each entity's name is sent to Wikidata. Requests identify themselves as "Scholaris/2" with the site's address. ## Your own keys You can add your own Gemini, OpenRouter, TypeSafe, Mistral, Voyage, Cohere, Jina or ZeroEntropy keys so your library uses your accounts. They are stored encrypted with AES-256-GCM, with a key derived per user, and are only used if you switch them on. ## Deleting - **Deleting a document** also deletes everything derived from it: pages, vectors, images. - **Deleting your search history**, or pausing it. - **Exporting everything** as a ZIP before you leave. - **Deleting your account:** your keys, settings, usage, invitations, notifications and shared links are deleted; your library is emptied; vectors and files are purged in the background. A row with your identifier remains, marked as deleted, with no email or name, so it cannot be reused. The exception: if someone copied a document from a library you shared with them, their copy remains theirs. ## Contact For anything about your data, [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- # Plans, coupons and limits URL: https://scholaris.joseluissaorin.com/en/knowledge/plans > A free plan to start, Pro for real work, and the home version with no quotas. How much fits in each, how pages and minutes are counted, and how coupons work. ## The plans | | Free | Pro | Home version | | --- | --- | --- | --- | | Documents | 25 | 5,000 | unlimited | | Pages or minutes read per month | 1,500 | 60,000 | unlimited | | Searches per day | 100 | 5,000 | unlimited | | Autocites per month | 5 | 500 | unlimited | | Storage | 1 GB | 100 GB | your disk | | Largest file | 200 MB | 4 GB | 16 GB | | API requests per minute | 120 | 600 | no practical limit | The free plan does not expire. The Pro price is shown in the app (Ajustes, Settings), where you subscribe; payments are handled by Clerk. The home version is free, but you pay the AI providers you use with your own keys. ## How it is counted - A page of a PDF, an EPUB or a photo counts as one page; a minute of audio or video counts as one page. - Monthly quotas reset on the 1st of each month and daily ones at midnight, UTC. - Uploading an identical file again does not count: it is recognised by its hash and the existing one is returned. ## Economy mode When filling a large library you can choose economy mode: hard pages are read through Gemini's batch API at half price, and take hours instead of seconds. On a 43-page scanned play, the cost fell by 46 %. ## Coupons A coupon is a code shaped like `SCHO-XXXX-XXXX` that grants a plan for a number of days or for good. Redeem it in Ajustes (Settings). A coupon never downgrades an account; if you already have that plan for a limited time, it extends it. To stop guessing, ten failed attempts per hour are allowed, and Scholaris only stores an encrypted hash of each code. If you think your class or your project should have one, write to [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). ## Trying it without an account The [demo](https://scholaris.joseluissaorin.com/?demostracion) opens the app with a sample library that lives in your browser: nothing is uploaded and no sign-up is needed. --- # What we measured, and under what conditions URL: https://scholaris.joseluissaorin.com/en/knowledge/performance > Reading times and costs, folio accuracy, search quality, invented citations and latency, as they came out of the benchmark of 6 October 2026, and what those figures do not say. ## Conditions Everything was measured on **6 October 2026**, from a Mac in Tenerife over a home connection, against the cloud APIs (Gemini 3.5 Flash-Lite and 3.8 Flash, Gemini Transcribe, Gemini Embedding 2 at 1536 dimensions, Jev, Crossref and OpenAlex) with no local GPU. **It was not measured on Cloudflare**, which should be more stable; nor was the latency of "ask". Dollar figures are provider costs, not prices. ## Reading a document | Document | Before (first Scholaris) | Now | Cost | | --- | --- | --- | --- | | Video interview, 54 min | 3 h 39 min | 45 s (30 s without speaker attribution) | $0.41 | | Video interview (Cortázar), 2 h 2 min | the 83-min ones took 5 h 43 min | 70 s | $0.94 | | *The Discarded Image*, 245 pp., with old OCR | not recorded | median 85 s over 16 runs (37 to 175 s) | $0.64 | | *El casamiento en la muerte*, 43 scanned pp., 17th century | not recorded | 27 s | $0.20 | | *Attention Is All You Need*, 15 digital pp. | not recorded | 12 s | $0.04 | | *El perseguidor*, 37 digital pp. | not recorded | 15 s | $0.04 | Times are end to end, conversion included. The first pages are searchable earlier: in under a tenth of a second for a digital PDF, in about 3 to 6 seconds for a scan and in about 3 to 6 seconds for a video. The slow runs of the 245-page book came from the API or the network, not the code. Very short media (one minute) now take a little longer to be ready than before: from 5 to 6 seconds it went to between 8 and 10. **Economy mode** cut the cost of the *Casamiento* from $0.206 to $0.111 (46 % less), in exchange for taking 6.9 minutes instead of 19.7 seconds. Transcription costs about $0.005 a minute. ## Reading well - **Character error rate (CER)** on seventeenth-century drama: 0.007 with Gemini 3.8 Flash, against 0.073 for the first Scholaris's reader. Measured on a single hand-transcribed page. - ***The Discarded Image***: the first Scholaris left 195 pages empty (74,000 characters in all); now there are 358,000 characters and 7 empty pages. - **Printed folio:** 64 of 64 exact on the pages checked by eye (45 of 45 with the number visible). In *The Discarded Image*, 229 folios read agree with the PDF's labels on 229 of 232 pages. - **Speakers:** attribution right in 20 of 21 turns and in 23 of 23 in two interviews (25 random turns each, annotated by reading the text). - **Records:** 55 of 55 checks right across 9 documents after checking against external sources (30 of 55 before). ## Searching well Over 183 queries and 8,879 relevance judgments. The judgments are "silver": two models and an arbiter made them, not people; they agreed 82 % exactly and 99 % within one point, and of 30 checked by hand, the author agreed with 29. | System | nDCG@10 | Recall@20 | MRR | | --- | --- | --- | --- | | First Scholaris | 0.695 | 0.618 | 0.925 | | Lexical only | 0.564 | | | | Visual only | 0.199 | | | | Semantic only | 0.804 | 0.860 | 0.931 | | Hybrid, not reranked | 0.818 | 0.884 | 0.942 | | **Current** (hybrid, reranked) | **0.889** | **0.904** | **0.979** | By query type (nDCG@10 of the current system): cross-language 0.897; literal 0.880; audio and video 0.882; old spelling 0.814; visual 0.950 (only 4 queries). **Old spelling:** on Lope de Vega's *El casamiento en la muerte*, with 30 queries, average recall per query went from 32 % to 100 % with the modernised-spelling layer; on modern documents, 250 of 250 searches returned exactly the same. **Latency:** median of about 650 ms from home, no caches (about 360 ms for the query vector and about 300 ms for the judge); without reranking, about 350 ms with nDCG 0.818. In the measured runs, median and 95th percentile were 687 and 856 ms. ## Citing well Over 41 claims (30 supported and 11 not): **95.1 % precision, 96.7 % recall, zero invented citations**, zero citations on unsupported claims and the verification verdict right 100 % of the time. Each claim took 998 ms and cost $0.0054. ## What these figures do not say - They come from a home connection, not Cloudflare, and do not include the latency of "ask". - 1,594 relevance judgments with an arbiter or a disagreement are still awaiting human review. - The CER was computed on a single page. - The benchmark is small and was made by the person who makes Scholaris. If you want to repeat it with your documents, write to us. --- # Frequently asked questions URL: https://scholaris.joseluissaorin.com/en/knowledge/faq > Whether it invents citations, whether it works with seventeenth-century scans or video interviews, what happens to your data, whether it works offline, and other questions, answered plainly. ## About citations ### Does it invent citations? No. No model writes a citation: author, year and page come from the document's record and from the anchor stored when it was read. When a model drafts an answer, it may only cite the passages it is given, and every note is checked against the text before you see it. If no passage supports something, the answer says so. In the 41-claim benchmark, invented citations were zero. See [Citations](https://scholaris.joseluissaorin.com/en/knowledge/citations.md). ### Is the page the PDF's or the book's? The book's: the number printed on paper, even when the PDF starts at the cover, the book uses roman numerals in the preface or skips plates. If the book prints no number, the physical position is cited in brackets, "p. [12]", instead of inventing one. See [Formats and anchors](https://scholaris.joseluissaorin.com/en/knowledge/formats.md). ### Which citation styles does it have? APA, Chicago (author-date and notes), MLA, Harvard, ISO 690 (author-date in Spanish and English, and numeric) and IEEE, in Spanish, English, French and Italian. It exports BibTeX, RIS and CSL-JSON, and inserts citations into a .docx with real Word footnotes. ## About what it reads ### Does it work with scanned seventeenth-century books? Yes, it is one of the cases it was built for. It reads pages with a vision model that keeps the original spelling ("dixo", "muger", u/v, ç) and leaves out running heads, catchwords and signatures; it recognises foliation by leaves ("23r", "23v"). So that those spellings show up when you search for "dijo" or "mujer", every passage carries a modernised-spelling layer used only for searching. On a 43-page play by Lope de Vega, the character error rate was 0.7 % and search recall went from 32 % to 100 %. See [Search](https://scholaris.joseluissaorin.com/en/knowledge/search.md#grafia). ### And with video interviews? Yes. It transcribes word by word, tells speakers apart and names them, cuts the text into 30-to-60-second stretches and cites the exact minute ("12:04-12:40"). Video frames are searchable too. A two-hour interview was ready in 70 seconds. YouTube is read from its address, without downloading the video. See [Audio and video](https://scholaris.joseluissaorin.com/en/knowledge/media-player.md). ### Does it read Latin? Yes. The search layer evens out u/v, i/j, æ/œ and the enclitics, and reduces words to their stem with the Schinke et al. stemmer, so "urbis" finds "Vrbs". ### Which formats does it take? PDF (digital and scanned), photos (JPG, PNG, HEIC…), EPUB, Word, ODT, RTF, HTML, Markdown, PowerPoint, Keynote, ODP, Excel, ODS, CSV, audio (MP3, WAV, M4A…), video (MP4, MOV, WebM…), web pages, YouTube, Vimeo and podcasts. The full list, with each one's anchor, is in [Formats and anchors](https://scholaris.joseluissaorin.com/en/knowledge/formats.md). ## About your data ### What happens to my data? It is stored on Cloudflare, in a space of its own for your account, and it is not used to train models. To read and search it, pages, audio and queries are sent to Google Gemini, Cloudflare Workers AI, OpenRouter and TypeSafe, each for a specific task; open bibliographic databases only receive titles, authors, ISBNs and DOIs. You can export everything and delete your account whenever you like. The detail, provider by provider, is in [Privacy](https://scholaris.joseluissaorin.com/en/knowledge/privacy.md). ### Can anyone else see my library? Only if you share it: by inviting someone by email or creating a public link, which you can protect with a password, set to expire or revoke. Those pages ask search engines not to index them. ### Can I use it offline? Partly. The home version keeps everything on your computer and, with InferBox, computes vectors, transcripts and answers there; but it still needs a cloud provider to read pages. What does work fully offline is opening and searching an .spdf that has already been read, with the Python SDK. See [Self-hosting](https://scholaris.joseluissaorin.com/en/knowledge/self-hosting.md). ### Is it open source? The Python SDK is (EUPL-1.2). The rest of the code will be published under the same licence when its review is finished; there is no date yet. ## About using it ### Does it write my paper? No, and it will not. It searches, points, cites and verifies; thinking and writing remain yours. The "cite" function gives you back your own text with the citations in place, not a new text. ### Does it search the internet for papers? No: it works on what you give it. To discover literature you do not have yet, use an academic search engine and bring what you find here. The citation graph does tell you which works your books cite that you do not have. ### How much does it cost? There is a free plan (25 documents, 1,500 pages or minutes a month, 100 searches a day) and a Pro plan, whose price is shown in the app. The home version has no quotas. See [Plans](https://scholaris.joseluissaorin.com/en/knowledge/plans.md). ### Can an agent or a program use it? Yes, through API v1 or the MCP server (Claude, Cursor and others). See [API v1 and the MCP server](https://scholaris.joseluissaorin.com/en/knowledge/api-and-mcp.md) and the [page for agents](https://scholaris.joseluissaorin.com/en/agents.md). ### Who makes it? José Luis Saorín Ferrer, philologist and programmer, in Santa Cruz de Tenerife. Write to [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- # Scholaris glossary URL: https://scholaris.joseluissaorin.com/en/knowledge/glossary > The words Scholaris uses, from folio and anchor to pliego, imprenta, SPDF or watcher, with what they mean here and, where it helps, in the tradition of the book. ## The words Anchor (ancla) : The exact place of a passage inside its document: the physical page and the printed folio, the second of a recording, the slide, the sheet and rows of a table, or the section and paragraph of a web page with its access date. Every citation is written from an anchor. Autocite (autocita) : The function that reads your draft, finds a passage in your library for each claim and proposes the citation with its page; you accept or discard. Notebook (cuaderno) : Markdown notes with cards of verified passages that can be checked again against the library. Vector space (espacio vectorial) : The model, version and dimensions used to compute some vectors, e.g. "gemini-embedding-2@1536". An .spdf can carry several. Record (ficha) : A document's metadata (authors, title, year, edition, publisher, DOI, ISBN…), each field with its provenance. Folio : In early books, each numbered leaf, with its recto (r) and verso (v): "fol. 23v". In Scholaris, by extension, the page number printed on paper, as opposed to the page's position in the file. Foliation (foliación) : How a book is numbered: by pages or leaves, in roman and arabic numerals, with jumps and unnumbered plates. Scholaris rebuilds it page by page. Fragment (fragmento) : The passage that is searched and cited: a stretch of text with its start and end anchors. Modernised spelling (grafía modernizada) : A copy of the text with spelling brought up to date ("dixo" → "dijo"), stored separately and used only for searching. The cited text always keeps its original spelling. Hash (huella) : The SHA-256 digest of a file; two files with the same hash are the same file. Imprenta (the press) : Scholaris's converter: it recognises each file's type and prepares it for reading (pages as images, text, audio, frames). It works in the browser and, for what the browser cannot do, in a server container. Jev : TypeSafe's model that acts as judge: it reranks results, decides whether a passage supports a claim and chooses between page-number candidates. Reader (lector) : The model that reads pages as images and returns their faithful text, notes, header and footer, and the folio it sees printed. Manicule (manícula) : The little hand with an outstretched finger that early readers drew in margins to point at a passage. It is Scholaris's emblem. Pliego : In printing, the large sheet that, folded, makes a gathering of the book. In Scholaris, the group of four pages sent at once to the reader. Provenance (procedencia) : Where each piece of data comes from: which source gave the record, which reader read the page, how confidently and who corrected it. Catchword (reclamo) : In early books, the word printed at the foot of a page that repeats the first word of the next. Scholaris keeps it out of the passage text. SPDF : Scholaris's open format for a document already read: a gzip-compressed SQLite database with text, anchors, figures, vectors and provenance. Current version: 4.1. .scholaris package : A whole library in a ZIP: one .spdf per document, a manifest and a rights notice. Stretch (tramo) : In audio and video, the citable unit: 30 to 60 seconds of transcript ending at a sentence boundary. Unit (unidad) : What is cited as a whole: a page, a stretch of audio, a slide, a chunk of a spreadsheet. Watcher (vigilante) : A saved search that repeats itself and tells you when what it finds changes. --- # Scholaris beside Zotero, Elicit, NotebookLM and the rest URL: https://scholaris.joseluissaorin.com/en/knowledge/alternatives > What each tool does, when another one suits you better than Scholaris, and what Scholaris does that the others do not set out to do. No figures about others that we cannot check. ## How to read this page We have not measured the other tools with our benchmark, so there are no figures of theirs here and no tables of ticked boxes. What follows describes what each one is designed for, as of this review, and how Scholaris differs. Tools change fast: if something we say about them is no longer true, tell us and we will fix it. ## Zotero A free, open-source reference manager: it captures references from the browser, stores PDFs, annotates them and formats citations in Word, LibreOffice and Google Docs with CSL styles. For collecting and organising references it is hard to beat, and Scholaris does not try to replace it. **When to choose Zotero:** to build and keep your bibliography, collaborate in reference groups and cite while you write in your word processor. **What Scholaris adds:** reading the content of those sources (old scans, audio and video included), finding the passage by meaning and returning the citation with the exact printed page. They get on well: Scholaris imports Zotero's BibTeX or RIS with its file folder and matches every entry to its document. ## Elicit A research assistant that searches the published scientific literature, summarises papers and extracts data from them into tables for systematic reviews. **When to choose Elicit:** to discover papers you do not have yet and compare many empirical studies at once. **What Scholaris adds:** it works on your own library, not on a corpus of papers: seventeenth-century scans, specific editions, recorded interviews, notes; and it cites the printed page or the second of your copy. ## NotebookLM Google's tool for talking with the sources you upload, with references to the source passages and audio overviews. **When to choose NotebookLM:** to get a quick sense of a small set of sources, free and with nothing to install, if you are happy for Google to process them inside its product. **What Scholaris adds:** citations in bibliographic styles (APA, Chicago, MLA, ISO 690…) with the printed folio, ready for a footnote; foliation of early books and a modernised-spelling layer; verification with temporal logic; an open format (SPDF) to take your read library with you; an API and MCP, and a version you install on your own computer. And it does not summarise books for you: that is not what it is for. ## ChatGPT or Claude with uploaded files General assistants read the files you attach to a conversation and answer about them, often very well. **When to choose them:** to talk, explore ideas or draft with help, over a handful of documents. **What Scholaris adds:** a persistent library of thousands of documents instead of one conversation's files; citations that come from the stored anchor rather than the model, checked before they reach you; the printed page and the second; and the rule of not writing for you. In fact they can be used together: Claude can query your Scholaris library over [MCP](https://scholaris.joseluissaorin.com/en/knowledge/api-and-mcp.md) and cite from it. ## Adobe Acrobat with AI Acrobat's assistant answers questions about PDFs and generates summaries inside the most widespread PDF reader. **When to choose it:** if you already work in Acrobat and asking about the PDFs you have open is enough. **What Scholaris adds:** many formats besides PDF (EPUB, audio, video, web pages, slides), search across the whole library at once, citation styles and printed pages, and the text of old scans read faithfully. ## In short | If you need to… | Look at… | | --- | --- | | Collect references and cite while writing in Word | Zotero (and, for the content, Scholaris beside it) | | Discover scientific papers you do not have yet | Elicit or an academic search engine | | Get a quick sense of a few sources | NotebookLM or a general assistant | | Find the exact passage in your library and cite it with its page or second | Scholaris | --- # Changelog URL: https://scholaris.joseluissaorin.com/en/knowledge/changelog > What has changed in Scholaris, from the first version to the second, told by milestones and in the order they were made. ## October 2026 - **6 October:** scholaris.joseluissaorin.com starts serving the second version of Scholaris, on Cloudflare. The first version is kept in reserve. - This knowledge base, with a Markdown twin of every page, [/llms.txt](https://scholaris.joseluissaorin.com/llms.txt), [/llms-full.txt](https://scholaris.joseluissaorin.com/llms-full.txt) and the [page for agents](https://scholaris.joseluissaorin.com/en/agents.md). - The Python SDK, ready to be published on PyPI as `scholaris-sdk`, under the EUPL-1.2. - Exporting an .spdf from the server, with a 24 MB cap on embedded files and a notice of what is left out. - Redoing a document's record when the evidence of its identity changes (another episode, another guest), with a dry-run mode. - Redoing only a document's figures, with the estimated cost shown first. ## The second version, in the order it was made The second version of Scholaris was written anew. Its milestones, first to last: 1. **Foundations:** the SPDF 4.0 format and the imprenta, the converter that reads any file in the browser. 2. **Search** by three fused routes (lexical, semantic and visual) and answers with validated citations; the official CSL styles. 3. **Reading pages in batches**, printed folios, records, fragments, vectors and figures; the concept map. 4. **Citations:** temporal logic, autocite with a judge, .docx insertion and BibTeX. 5. **The API on Cloudflare Workers**, with Clerk sign-in and encrypted own keys; SPDF on Workers and migration of old SPDFs. 6. **The web app** (library, reader, search, writing, exploring), the **home version** on Node and SQLite, the **desktop executable** and the Docker image. 7. **A new reader**, Gemini 3.8 Flash, chosen with the benchmark. 8. **Media:** YouTube by address, podcasts, Vimeo and direct links. 9. **The MCP server with OAuth 2.1** and migration of first-version accounts. 10. **Reading in three stages** (readable, searchable, vectorised), **economy mode** and the **entity graph** linked to Wikidata. 11. **The quality benchmark:** nDCG@10 of 0.889 and no invented citations. 12. **Two-stage search** and the **new player**, highlighting the word being spoken. 13. **API v1** with its guide at [/en/api](https://scholaris.joseluissaorin.com/en/api.md), the multi-user home version and shared libraries. 14. **Drawings and motion** in the app; records for radio and TV programmes; OpenAlex with a cache. 15. **Getting ready for production.** ## The first version The first Scholaris ran on a home server and stored documents as SPDF 1 to 3. Its SPDFs still open: they are migrated to format 4 on the fly. Compared with it, the second reads a 54-minute interview in 45 seconds instead of three hours and 39 minutes, and finds better (nDCG@10 from 0.695 to 0.889). See [Performance](https://scholaris.joseluissaorin.com/en/knowledge/performance.md). --- # For agents: how to read Scholaris and how to use it on someone’s behalf URL: https://scholaris.joseluissaorin.com/en/agents > What an agent can read on this site and in which format, how to act on a person’s library through API v1 or MCP, with which permissions and limits, and the rules for citing without inventing. ## If you only want to understand what Scholaris is Everything public is built to be read without JavaScript and without scraping HTML. Every page has a Markdown twin at the same address ending in `.md`, and is also served as Markdown if you ask with the header `Accept: text/markdown`. Pages also carry schema.org JSON-LD (SoftwareApplication, TechArticle, FAQPage, HowTo, DefinedTermSet). ```sh # Any public page, as Markdown curl -s https://scholaris.joseluissaorin.com/en/knowledge/formats.md curl -s -H "Accept: text/markdown" https://scholaris.joseluissaorin.com/en/knowledge/formats # Everything at once curl -s https://scholaris.joseluissaorin.com/llms.txt curl -s https://scholaris.joseluissaorin.com/llms-full.txt ``` ## Files for machines | Address | What it is | | --- | --- | | [/llms.txt](https://scholaris.joseluissaorin.com/llms.txt) | Short index with a link to every Markdown page | | [/llms-full.txt](https://scholaris.joseluissaorin.com/llms-full.txt) | The full text of every public page, Spanish and English | | /any/page.md | The Markdown twin of each public page | | [/api/v1/llms.txt](https://scholaris.joseluissaorin.com/api/v1/llms.txt) | How to use the API without inventing citations | | [/api/v1/openapi.json](https://scholaris.joseluissaorin.com/api/v1/openapi.json) | OpenAPI 3.1 specification of API v1 | | [/.well-known/api-catalog](https://scholaris.joseluissaorin.com/.well-known/api-catalog) | API catalogue (RFC 9727) | | [/.well-known/mcp/server-card.json](https://scholaris.joseluissaorin.com/.well-known/mcp/server-card.json) | MCP server card | | [/.well-known/oauth-protected-resource/mcp](https://scholaris.joseluissaorin.com/.well-known/oauth-protected-resource/mcp) | OAuth metadata of the MCP resource | | [/.well-known/oauth-authorization-server](https://scholaris.joseluissaorin.com/.well-known/oauth-authorization-server) | OAuth authorisation server metadata | | [/.well-known/security.txt](https://scholaris.joseluissaorin.com/.well-known/security.txt) | Where to report a security problem | | [/sitemap.xml](https://scholaris.joseluissaorin.com/sitemap.xml) | Every public page, with its date and languages | | [/robots.txt](https://scholaris.joseluissaorin.com/robots.txt) | What may be crawled | ## What may be crawled The public pages (the front page, this knowledge base, the API guide and this page) may be read, indexed, quoted and used to answer, including for training models: [/robots.txt](https://scholaris.joseluissaorin.com/robots.txt) says so, with a line for each known crawler (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, CCBot and others). The app, the API and anyone's libraries are **not** to be crawled: the public links people share explicitly ask not to be indexed, and their content belongs to them. ## If you act on a person's behalf You need that person to give you access; there is no anonymous access to any library. There are two ways: 1. **MCP with OAuth.** If your client speaks MCP, connect to `https://scholaris.joseluissaorin.com/mcp`. You register yourself (dynamic client registration), the person signs in to Scholaris and grants you read-only access. Tools: `search`, `cite`, `open_page` and `verify_claim`. 2. **API v1 with a key.** The person creates a key in Ajustes (Settings) → Claves de API (API keys) and gives it to you. With the `lectura` (read) scope you can search, read, ask and verify; with `escritura` (write), also upload, delete and cite a whole text; with `mcp`, use it with the MCP server. Base: `https://scholaris.joseluissaorin.com/api/v1`. ```sh curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=la música me metía en el tiempo" -d k=5 -d formato=markdown ``` All of it is explained in [API v1 and the MCP server](https://scholaris.joseluissaorin.com/en/knowledge/api-and-mcp.md) and, with a recorded session, in the [API guide](https://scholaris.joseluissaorin.com/en/api.md). ## Limits you must respect - **Rate:** 120 requests per minute on the free plan and 600 on Pro. On a 429, wait the seconds in `Retry-After`. - **The person's quotas:** searches per day, pages or minutes per month, autocites per month (see [Plans](https://scholaris.joseluissaorin.com/en/knowledge/plans.md)). A 402 means they are used up: tell them, do not keep trying. - **Files:** up to 95 MB per request on API v1. - **Retries:** send `Idempotency-Key` when uploading and citing so nothing is duplicated. - **What takes time:** uploading and citing wait by default; not to block, use `esperar=0` or `Prefer: respond-async` and poll `progreso_url`. ## Rules for citing 1. Cite only what Scholaris returns, and copy `cita` (citation) and `localizador` (locator) verbatim. 2. When you quote, copy the literal `texto`; when you paraphrase, still attach the citation. 3. Give the `enlace` (link): it is how the person checks the page or the second. 4. If the search finds nothing, say so. Do not fill the gap from memory or invent a page. 5. Do not present an answer from `preguntar` (ask) as your own: its footnotes are the important part. 6. Do not write the person's work and pass it off as theirs. Scholaris exists so that they find and think. ## Errors Every error response has a stable `codigo` (code), a `mensaje` in Spanish, a `message` in English and a link to the documentation. API fields are in Spanish and lower case. ## Contact To report a bug, ask for a higher rate or propose an integration: [jl@joseluissaorin.com](mailto:jl@joseluissaorin.com). --- URL: https://scholaris.joseluissaorin.com/en # A hand that points to the page, then steps aside. > Scholaris reads your books, your papers, your notes and your interviews, and when you ask it something it doesn’t hand you an opinion. It points to the passage, the printed page, the exact minute something was said, so that the thinking stays yours. This is Scholaris's front page, written as an essay in eight chapters. The technical details (formats, citations, privacy, API, performance) are in the knowledge base: https://scholaris.joseluissaorin.com/en/knowledge.md ## Chapter the first, which treats of abundance, and of how little is found in it Every student has a folder called “to read”, which later became “to read, really” and later still “thesis_final_2”, and somewhere inside it is the paragraph you need. You underlined it one afternoon. You know it was on a left-hand page, somewhere near the end of a chapter, with a pencilled note of yours beside it that said something like “yes!”. You cannot find it. There has never been so much to read, nor has keeping it ever been so easy, and yet the commonest experience of anyone who studies is still this one: knowing that something exists and being unable to reach it. Our libraries fit in a pocket and behave like that other one, the endless library of hexagonal galleries, which holds every possible book and opens none of them at the right page. Abundance did not solve the problem of finding; it enlarged it until it became invisible, because now the fault seems to be ours. > Cf. J. L. Borges, “The Library of Babel”, in *Ficciones* (1944). No page number: this page doesn’t cite what it hasn’t checked either. ## Chapter the second, which treats of the archive, and of who decides what can be said An archive is never innocent. What is kept, the order in which it is kept and what is left out decide, before anyone else does, what can later be thought. For centuries copyists, librarians and indexes made that decision; today it is made, more and more, by machines that do not show their shelves. **When you cannot tell where what you read comes from, you cannot argue with it, and whatever cannot be argued with ends up in charge.** That is why provenance matters so much to us. It is no academic ornament, no footnote that nobody reads, but the least a reading needs in order to be yours: knowing who said it, where, in which edition, on which page and with what degree of certainty, and being able to go and check. > Cf. M. Foucault, *The Archaeology of Knowledge* (1969), “The historical *a priori* and the archive”. ## Chapter the third, which treats of the machines that answer A search engine gives you ten links and a page of advertisements; a chatbot gives you a rounded answer, sure of itself, without a single page number. The first leaves you at the door of somebody else’s building; the second tells you what is inside without letting you in and, now and then, invents the books, the authors and the pages in the same calm voice it uses when it is right. Both of them flatten. They turn the difference between one text and another (between a 1925 edition and a 1970 translation, between what someone said and what someone is said to have said) into a smooth surface where everything weighs the same. And both, without meaning to, take away the work you most wanted to do yourself: read the passage, distrust it, set it beside another one, and think. ## Chapter the fourth, which treats of the pointing hand In the margins of old books a small hand with an outstretched finger appears again and again. Palaeographers call it a manicule, and it was for exactly this: to tell some future reader “look here”, without telling them what to think of what they were about to see. Scholaris wants to be that hand. Ask it something and it gives you back the passage, the book and the printed page (the real one, the one you would put in a footnote) or the exact minute of the interview in which someone said it. If it cannot find it, it says so. **It does not invent citations:** each one is checked against the text you gave it before you ever see it, and every piece of data carries its provenance in plain view, where it came from, how sure we are of it and who corrected it. > «En un lugar de la Mancha, de cuyo nombre no quiero acordarme, no ha mucho tiempo que vivía un hidalgo de los de lanza en astillero…» > Miguel de Cervantes, *El ingenioso hidalgo don Quijote de la Mancha*, Madrid, Juan de la Cuesta, 1605, fol. 1r And if what you have is a recording, minute 12:04, word for word. ## Chapter the fifth, which treats of how an ordered library begins to think with you A well-read library reveals relations nobody had written down: two authors who never cite each other and talk about the same thing, a concept that changes its name when it crosses a border, a question that comes back every forty years in different clothes. Scholaris reads your whole library at once and draws those relations (people, works, places, concepts, dates) like a chart of the sky, with constellations that were in no single book and only appear when the books are put together. The machine draws the lines, always anchored to the page they come from; what they mean, which of them matter and where they lead is for you to decide. We want a structure that multiplies thought instead of closing it, one that works less like a tree, with its trunk and its hierarchy, and more like a rhizome, where any point can connect to any other and every connection is the beginning of an idea that does not exist yet. > Cf. G. Deleuze and F. Guattari, *A Thousand Plateaus* (1980), “Introduction: Rhizome”. ## Chapter the sixth, which treats of what Scholaris will never do It will not write your essay. It will not summarise a book so that you can skip it. It will not decide for you which quotation is the good one or which author is right. Thinking, judging and writing remain yours, and they do as a matter of principle: a thought that has not passed through the body of the person who signs it is too much like a rumour. ## Chapter the seventh, which treats of what it does, told quickly - **Anything goes in** (PDF · EPUB · DOCX · MP3 · MP4). Digital or scanned PDFs, photos of pages, EPUB, Word, slides, audio, video and links: all of it becomes a library you can search and cite. - **The printed page** (p. 145 · xiv · fol. 1r). The real folio, the one you would put in a footnote, even when the book starts in roman numerals or the pagination jumps. - **The exact minute** (12:04). Interviews, lectures and recordings transcribed word for word; every citation jumps to its second. - **Questions to the whole library** (verified citations). In any language, Old Spanish and Latin included, with answers whose citations are checked one by one before they reach you. - **Your citations, your style** (CSL · BibTeX · RIS). APA, MLA, Chicago or whatever your journal asks for; BibTeX and RIS; and an autocite that reads your draft and proposes each reference with its page. - **On your computer or in the cloud** (local · cloud). Install it at home, with no account and your library on your own disk (reading uses your own AI keys or your own inference server), or use it in the cloud from anywhere. - **An open format that is yours** (.spdf). Your whole library, already read, in an open SPDF file you can take with you whenever you like. ## Chapter the eighth and last, which treats of whom it is for For those who read in order to write. For the doctoral student with three hundred PDFs and a supervisor who wants the page number; for the lecturer preparing a class with twenty years of underlining; for the journalist going back to a two-year-old interview in search of one sentence; for the translator, the archivist, the first-year student who doesn’t yet know they will need it, and for anyone who has ever said “I read it somewhere”. ## Here ends the front page Scholaris is made by José Luis Saorín Ferrer, philologist, programmer and artist, in Santa Cruz de Tenerife. The drawings on this page are written stroke by stroke, as if traced, and a program gives them a pulse: they always tremble the same way, but they tremble. *Santa Cruz de Tenerife, autumn 2026* Start: https://scholaris.joseluissaorin.com/?entrar · Try it without an account: https://scholaris.joseluissaorin.com/?demostracion --- URL: https://scholaris.joseluissaorin.com/en/api # Upload anything, search, ask and cite. > One header, nine verbs, JSON in and out (or Markdown, which agents read better). Every passage comes with its exact printed page or second and a link to the reader, and no citation comes from a model: they all come from the stored anchor. Base URL: `https://scholaris.joseluissaorin.com/api/v1` · OpenAPI: https://scholaris.joseluissaorin.com/api/v1/openapi.json · llms.txt: https://scholaris.joseluissaorin.com/api/v1/llms.txt ## Start in a minute 1. Create a key in Ajustes (Settings) → Claves de API (API keys), with the read and write scopes. It is shown only once: keep it. 2. Upload a file. The request waits until it has been read (up to a minute; if it takes longer, it answers 202 with where to ask). 3. Search. Every passage brings a citation ready to paste, the exact locator and the link to the reader at that page. ```sh export SCHOLARIS=sch_… # tu clave curl -s https://scholaris.joseluissaorin.com/api/v1/documentos -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/pdf" --data-binary @articulo.pdf curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=atención escalada" -d k=3 -d formato=markdown ``` The examples can be copied and pasted as they are: you only need the SCHOLARIS variable with your key. ## The nine verbs Everything hangs from one base address. Field names are in Spanish and lower case; this guide translates them. ### POST /api/v1/documentos Upload anything: the file as the body (PDF, EPUB, DOCX, slides, spreadsheets, audio, video, images), a multipart form with the field `archivo` (file), or JSON with `url` (a web page, a PDF, YouTube, Vimeo, a podcast). If the file was already there (same SHA-256), you get the existing one with `duplicado` (duplicate). ```sh # Un fichero, como cuerpo (con su tipo y su nombre) curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos?nombre=libro.epub" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/epub+zip" --data-binary @libro.epub # Una URL (web, PDF, YouTube, Vimeo, pódcast), sin esperar curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos?esperar=0" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/json" -d '{"url": "https://www.youtube.com/watch?v=fNk_zzaMoSs"}' # Un formulario, con la ficha curl -s https://scholaris.joseluissaorin.com/api/v1/documentos -H "Authorization: Bearer $SCHOLARIS" \ -F archivo=@entrevista.mp3 -F titulo="A fondo" -F autores="Cortázar, Julio" -F anio=1977 ``` ### GET /api/v1/documentos List the library. Filter with `q` (title or author) and `estado` (status); page with `cursor`. ```sh curl -sG https://scholaris.joseluissaorin.com/api/v1/documentos -H "Authorization: Bearer $SCHOLARIS" -d estado=listo -d formato=markdown ``` ### GET /api/v1/documentos/{id} The record: title, authors, year, status, progress while it is processed and the full reference when it is ready. With `esperar=30` (wait) it long-polls until done. ```sh curl -s "https://scholaris.joseluissaorin.com/api/v1/documentos/ID?esperar=30" -H "Authorization: Bearer $SCHOLARIS" ``` ### DELETE /api/v1/documentos/{id} Delete it with everything derived from it (pages, vectors, images). ```sh curl -s -X DELETE https://scholaris.joseluissaorin.com/api/v1/documentos/ID -H "Authorization: Bearer $SCHOLARIS" ``` ### GET /api/v1/documentos/{id}/texto Read the text by pages or by a time range. `desde` (from) and `hasta` (to) take the printed page (`23`, `xiv`), the physical position in brackets (`[12]`) and, in audio and video, times (`1:06:56`). ```sh # Las páginas 23 a 25 tal como están impresas curl -sG https://scholaris.joseluissaorin.com/api/v1/documentos/ID/texto -H "Authorization: Bearer $SCHOLARIS" -d desde=23 -d hasta=25 -d formato=markdown # Un tramo de una entrevista curl -sG https://scholaris.joseluissaorin.com/api/v1/documentos/ID/texto -H "Authorization: Bearer $SCHOLARIS" -d desde=1:06:00 -d hasta=1:08:00 ``` ### GET /api/v1/buscar Search passages (hybrid search, reranked). Quote a phrase for a literal match. Each passage brings `cita` (citation), `localizador` (locator), `ancla` (anchor), `enlace` (link) and the literal `texto`. ```sh curl -sG https://scholaris.joseluissaorin.com/api/v1/buscar -H "Authorization: Bearer $SCHOLARIS" \ --data-urlencode "q=la música me metía en el tiempo" -d k=5 ``` ### POST /api/v1/preguntar Answer in Markdown with footnotes [^n]. The writer may only cite the passages it is given, and every note is checked. With `stream` it arrives as events. ```sh curl -s https://scholaris.joseluissaorin.com/api/v1/preguntar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"pregunta": "¿Qué le pasa a Johnny con el tiempo cuando toca?"}' ``` ### POST /api/v1/citar Your text back with the citations inserted (exact page, any CSL style) and the bibliography. It also takes a .docx and returns it cited, formatting intact. ```sh curl -s https://scholaris.joseluissaorin.com/api/v1/citar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"texto": "La música no saca a Johnny del tiempo: lo mete en otro.", "estilo": "chicago-author-date"}' # Un .docx, devuelto citado curl -s "https://scholaris.joseluissaorin.com/api/v1/citar?estilo=apa" -H "Authorization: Bearer $SCHOLARIS" \ -H "Content-Type: application/vnd.openxmlformats-officedocument.wordprocessingml.document" \ -H "Accept: application/vnd.openxmlformats-officedocument.wordprocessingml.document" \ --data-binary @trabajo.docx -o trabajo-citado.docx ``` ### POST /api/v1/verificar Whether your library supports a claim: `respaldada` (supported), `parcial`, `sin_respaldo` (unsupported) or `contradicha` (contradicted), with the probability and the passages. ```sh curl -s https://scholaris.joseluissaorin.com/api/v1/verificar -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" \ -d '{"afirmacion": "Johnny Carter dice que la música lo mete en el tiempo."}' ``` Every read accepts `formato=markdown` (or the header Accept: text/markdown). What takes time (uploading and citing) waits by default and answers with the result. Not to wait, `esperar=0` or the header `Prefer: respond-async`: the answer is a 202 with `progreso_url`. ## A real session Five commands recorded as they ran against the home version, with the benchmark files: a short story by Cortázar in PDF and a lecture in audio already uploaded. Long answers are cut where it says […]. ### Upload a PDF and wait until it is read ```sh curl -s "$B/documentos?esperar=120" -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/pdf" --data-binary @cortazar1959perseguidor.pdf | jq '{id, titulo, autores, estado, unidades, referencia}' ``` ```text { "id": "dmuwzohwaoqul34wy", "titulo": "El perseguidor", "autores": [ "Cortázar, Julio" ], "estado": "listo", "unidades": 37, "referencia": "Cortázar, J. (1996). El perseguidor." } ``` ### Search, in Markdown ```sh curl -sG "$B/buscar" -H "Authorization: Bearer $SCHOLARIS" --data-urlencode "q=la música me metía en el tiempo" -d k=2 -d formato=markdown ``` ```text # la música me metía en el tiempo ## 1. (Cortázar, 1996, p. 5) *El perseguidor* · p. 5 · [abrir](http://localhost:8795/lector/dmuwzohwaoqul34wy?u=5&f=dmuwzohwaoqul34wy%3At0.8) · `dmuwzohwaoqul34wy:t0.8` > Vaya si lo he oído; vaya si he tratado de escribirlo bien y verídicamente en mi biografía de Johnny. > > —Por eso en casa el tiempo no acababa nunca, sabes. De pelea en pelea, casi sin comer. Y para colmo la religión, ah, eso no te lo puedes imaginar. Cuando el maestro me consiguió un saxo que te hubieras muerto de risa si lo ves, entonces creo que me di cuenta en seguida. La música me sacaba del tiempo, aunque no es más que una manera de decirlo. Si quieres saber lo que realmente siento, yo creo que la música me metía en el tiempo. Pero entonces hay que creer que este tiempo no tiene nada que ver con... bueno, con nosotros, por decirlo así. > […] ``` ### An audio passage: the locator is the second ```sh curl -sG "$B/buscar" -H "Authorization: Bearer $SCHOLARIS" --data-urlencode "q=vectors as arrows in space" -d k=1 | jq '.pasajes[0] | {cita, localizador, enlace}' ``` ```text { "cita": "(3Blue1Brown, s. f., 0:11)", "localizador": "0:11", "enlace": "http://localhost:8795/lector/dmuwzm8ascxmk4ayv?t=11.3&f=dmuwzm8ascxmk4ayv%3Ac9tp39p" } ``` ### Ask ```sh curl -s "$B/preguntar?formato=markdown" -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" -d '{"pregunta":"¿Qué le pasa a Johnny con el tiempo cuando toca?"}' ``` ```text Cuando Johnny toca, experimenta que el tiempo cambia[^1][^2]. Aunque en un principio expresa que «la música me sacaba del tiempo», matiza que en realidad siente que lo «metía en el tiempo», entendiendo que se trata de un tiempo distinto que no tiene que ver con el orden cotidiano[^3]. Niega que se trate de una abstracción; más bien describe que simplemente cambia de lugar, comparando la experiencia con entrar en «un ascensor de tiempo» donde transcurren decenas de pisos en lo que dura una sola frase, mientras asuntos como las hipotecas o la religión dejan de existir para él[^4]. Asimismo, percibe la elasticidad del tiempo mediante la metáfora de una valija en la que, al tocar, es capaz de meter «cientos y cientos de trajes» o «una tienda entera»[^5]. Durante esos momentos musicales, Johnny siente que se abre una puerta y que vuela hacia un estado de seguridad donde queda «sobre todo sin tiempo, sin que después... sin que hubiera después... Por un rato no hubo más que siempre...»[^6]. Sin embargo, esa vivencia dura solo mientras se encuentra sumergido en la música; al dejar de tocar, cae de nuevo «de cabeza» en sí mismo[^6]. Bruno observa que en su ejecución Johnny «siempre está tocando mañana», adelantándose sin esfuerzo al hoy[^3], y viviendo intensamente lo que describe como su «cuarto de hora de minuto y medio», una alteración temporal frente al límite impuesto por los relojes[^1][^7]. [^1]: Cortázar, *El perseguidor* (1996), p. 9. [^2]: Cortázar, *El perseguidor* (1996), p. 5. [^3]: Cortázar, *El perseguidor* (1996), p. 5. [^4]: Cortázar, *El perseguidor* (1996), pp. 5-6. [^5]: Cortázar, *El perseguidor* (1996), p. 6. [^6]: Cortázar, *El perseguidor* (1996), p. 34. [^7]: Cortázar, *El perseguidor* (1996), pp. 16-17. --- Fuentes (confianza alta): […] ``` ### Cite your own text ```sh curl -s "$B/citar?formato=markdown" -H "Authorization: Bearer $SCHOLARIS" -H "Content-Type: application/json" -d '{"texto":"Johnny Carter siente que la música no lo saca del tiempo, sino que lo mete en otro. Ese tiempo no se parece al de los relojes.","estilo":"apa"}' ``` ```text Johnny Carter siente que la música no lo saca del tiempo, sino que lo mete en otro (Cortázar, 1996, p. 5). Ese tiempo no se parece al de los relojes (Cortázar, 1996, pp. 8-9). ## Referencias - Cortázar, J. (1996). El perseguidor. ``` ## From JavaScript With fetch and nothing else (Node 20 or the browser). Ten lines: upload, search and ask. ```js import { readFile } from 'node:fs/promises'; const B = 'https://scholaris.joseluissaorin.com/api/v1'; const h = { Authorization: `Bearer ${process.env.SCHOLARIS}` }; const pedir = async (ruta, o = {}) => (await fetch(B + ruta, { ...o, headers: { ...h, ...o.headers } })).json(); const doc = await pedir('/documentos?nombre=articulo.pdf', { method: 'POST', headers: { 'Content-Type': 'application/pdf' }, body: await readFile('articulo.pdf') }); const { pasajes } = await pedir(`/buscar?k=3&q=${encodeURIComponent('atención escalada')}`); for (const p of pasajes) console.log(p.cita, p.texto.slice(0, 80), p.enlace); const r = await pedir('/preguntar', { method: 'POST', headers: { 'Content-Type': 'application/json' }, body: JSON.stringify({ pregunta: '¿Qué es la atención multicabeza?' }) }); console.log(r.respuesta); ``` ## From Python The SDK has a five-verb facade over this same API. Install it with pip (it only needs requests): ```sh pip install scholaris-sdk ``` ```py from scholaris.api import Scholaris s = Scholaris("sch_…") # o la variable SCHOLARIS_CLAVE doc = s.subir("articulo.pdf") # también una URL for p in s.buscar("atención escalada", k=3): print(p["cita"], p["texto"][:80], p["enlace"]) print(s.preguntar("¿Qué es la atención multicabeza?")["respuesta"]) print(s.citar("La atención sustituye a la recurrencia.")["texto"]) print(s.verificar("El Transformer prescinde de la recurrencia.")["veredicto"]) ``` ## For agents There are two ways in: the MCP server, for agents that speak MCP, and this API, which any agent with an HTTP tool can use. In Claude Code, with a key that has the `mcp` scope: ```sh claude mcp add --transport http scholaris https://scholaris.joseluissaorin.com/mcp \ --header "Authorization: Bearer sch_…" ``` In Claude (web or desktop): Settings → Connectors → Add custom connector, with this address. You sign in to Scholaris and grant access; no key needed. ```text https://scholaris.joseluissaorin.com/mcp ``` In Cursor, Windsurf and other MCP clients: ```json { "mcpServers": { "scholaris": { "url": "https://scholaris.joseluissaorin.com/mcp", "headers": { "Authorization": "Bearer sch_…" } } } } ``` An agent that can only read the web finds its instructions in /llms.txt: the verbs, the fields and the rules to cite without inventing. ```sh curl -s https://scholaris.joseluissaorin.com/llms.txt ``` 1. Cite only what the API returns, and copy `cita` and `localizador` verbatim. 2. When you quote, copy the literal `texto`; when you paraphrase, still attach the citation. 3. Give the `enlace`: it is how the reader checks the page or the second. 4. If the search finds nothing, say so; do not fill the gap from memory. ## When something fails Errors say what happened in Spanish (`mensaje`) and in English (`message`), with a stable code for your program and a link to this guide: | Code | HTTP | Meaning | | --- | --- | --- | | `no_autenticado` | 401 | The key is missing, wrong or revoked. | | `prohibido` | 403 | The key lacks the scope (writing needs `escritura`). | | `peticion_invalida` | 400 | A field is missing or invalid: the message says which. | | `no_encontrado` | 404 | It does not exist or it is not yours. | | `conflicto` | 409 | It clashes with the current state; for example, the document has no text yet. | | `demasiado_grande` | 413 | This endpoint takes files up to 95 MB; for more, the SDK uploads in parts. | | `cuota_superada · requiere_pro` | 402 | A quota of the plan is used up. | | `limite_de_ritmo` | 429 | Too many requests: wait the seconds in Retry-After. | | `proveedor_fallo` | 502 | An AI provider failed: try again. | Rate limits and quotas are those of your plan, the same as in the app. To retry without duplicating, send an Idempotency-Key header when uploading and citing: with the same key, for 24 hours, you get the same resource back. ## Where every citation comes from When it reads a document, Scholaris stores an anchor for every passage: the physical page and the printed folio you see on paper, the second of an audio or a video, the slide, the section and paragraph of a web page. That is the `ancla` field. The `cita` and the `localizador` are written from that anchor, never from what a model says; the `enlace` opens the reader at that same page or second. If a passage is not in your library, the API will not cite it.