---
title: "Scholaris on your computer, or in your own Cloudflare account"
description: "The same app with SQLite and your disk: on Node, in Docker, as a desktop executable or deployed to your Cloudflare account. What stays on your machine and how it works offline."
url: https://scholaris.joseluissaorin.com/en/knowledge/self-hosting
markdown: https://scholaris.joseluissaorin.com/en/knowledge/self-hosting.md
lang: en
alternate_es: https://scholaris.joseluissaorin.com/saber/version-local.md
updated: 2026-10-07
author: José Luis Saorín Ferrer (https://joseluissaorin.com)
---

# Scholaris on your computer, or in your own Cloudflare account

> The same app with SQLite and your disk: on Node, in Docker, as a desktop executable or deployed to your Cloudflare account. What stays on your machine and how it works offline.

## Four ways

The instance at [scholaris.joseluissaorin.com](https://scholaris.joseluissaorin.com/en.md) is the hosted version, the one paid for with the Pro plan. The same code also runs on your computer:

| How | What it is | Where your data lives |
| --- | --- | --- |
| Node | The same API on Node, SQLite (with sqlite-vec for vectors) and the disk, with a queue that resumes if interrupted; listens on port 8790 | The folder you choose |
| Docker | The same in a container with ffmpeg; a variant runs it offline with its models (Ollama, EmbeddingGemma 2 and Whisper), with or without an NVIDIA GPU | A Docker volume |
| Desktop | A single executable (Bun) for macOS, Windows and Linux, with the web app inside, that opens the browser on start | ~/Scholaris |
| Your Cloudflare account | A script creates the database, storage, vector index and queue, and deploys the Worker (needs paid Workers) | Your account |

The home version has no quotas: the "local" plan does not limit documents, pages or searches, and takes files up to 16 GB. It can have a single user, several without an external account, or use Clerk to sign in.

## Open source

The code will be published under the **EUPL-1.2** at [github.com/joseluissaorin/scholaris-v2](https://github.com/joseluissaorin/scholaris-v2). While the review is finished the repository is private: we are not giving a date. The exact installation commands will be in its README, which takes precedence over this page. The Python SDK already carries the same licence.

## What stays on your machine and what does not

In the home version, **your files, your library, the vectors and the index live on your disk**. You choose where the intelligence runs:

- **With your keys** (Gemini; OpenRouter, TypeSafe and Workers AI optional): the fastest and most faithful. Pages, audio and passages go to those providers, under their terms.
- **Offline** (`SCHOLARIS_SIN_CONEXION=1`): no cloud key at all. Everything runs on your machine or your network and nothing goes out to the internet; a network guard blocks and logs any attempt, and an end-to-end test checks it.

| Piece | Offline | With a cloud provider |
| --- | --- | --- |
| Reading pages (scans, photos, slides) | Your own vision model: Qwen3-VL 8B Instruct on Ollama, llama.cpp or vLLM | Gemini, OpenRouter (Mistral OCR) or Workers AI |
| Vectors | EmbeddingGemma 2 (text, image, audio and video, 768 dimensions), on its own server or on InferBox | Gemini Embedding 2 |
| Reranking | bge-reranker-v2-m3 | Jev (TypeSafe) or Workers AI |
| Judging citations and drafting | The same vision model, with probabilities from logprobs | Jev, Gemini or OpenRouter |
| Transcribing | Whisper large-v3-turbo (whisper.cpp) or Parakeet on InferBox | Gemini Transcribe or Whisper on Workers AI |
| Bibliographic records | Only what the document says (online catalogues open with `SCHOLARIS_CATALOGOS=1`) | Crossref, OpenAlex, Open Library, Wikidata |

What it costs, measured on an M4 Max Mac: an eighteenth-century scanned page takes about 26 s (the cloud reads the whole book in 13-20 s), with a CER of 0.06 on the hand-transcribed page against 0.006 with Gemini; 19 minutes of audio are ready in 2 min 34 s; search scores 0.804 nDCG@10 against 0.899; and every citation it accepts is correct and none is invented, though it leaves more claims without a citation than the cloud does (61.5 % recall against 100 %). On CPU alone, reading scans takes minutes per page. The full report, with how to reproduce it, is in the repository (`packages/proveedores/SIN-CONEXION.md`), and the all-in-one Docker setup in `deploy/docker/compose.sin-conexion.yml`.

On top of that, **opening and searching an .spdf that has already been read** works with none of the above, through the Python SDK (see [The SPDF format](https://scholaris.joseluissaorin.com/en/knowledge/spdf.md)).

## The same from outside

The home version speaks the same API (v1 and v2) and the same MCP as the cloud, so the Python SDK, the examples in the [API guide](https://scholaris.joseluissaorin.com/en/api.md) and agents work the same pointed at `http://localhost:8790`.
