Skip to main content
Ragen is self-hosted, so everything it stores stays where you install it. Storage is only half the question, though. A deployment can keep every document in Paris and still send prompts to a model provider outside the EU. This page shows how to configure every layer so that nothing is processed by a provider outside EU jurisdiction. It builds on Local Models, which covers the fully offline variant.
Jurisdiction, not just region. An EU region of a non-EU cloud provider keeps data physically in Europe, but the provider remains subject to the laws of its home country. For many organisations in the public sector, finance and healthcare, that is the deciding factor. This page is about the provider, not only the data centre.

The Stack, Layer by Layer

Where a model’s weights come from does not decide where your data goes. An open model developed anywhere, served on EU infrastructure or on your own GPU, processes your data only there.

Configuration

1

Pick EU hosting and storage

Run the stack on an EU provider or your own servers. For object storage, the quickest route is the bundled RustFS: answer RustFS to the storage question in create-ragen-app. To use a managed EU bucket instead, follow Storage, for example Scaleway Object Storage in fr-par, nl-ams or pl-waw.
2

Connect Scaleway Generative APIs

The shipped route table (infra/llm-gateway/routes.yaml) already carries the Scaleway routes: mistral-small-3.2, gpt-oss-120b, bge-multilingual-gemma2 and qwen3-embedding-8b. They only need credentials:
.env.local
Create the key in the Scaleway Console with the GenerativeApisFullAccess scope.Choosing Scaleway in create-ragen-app writes these for you.
3

Point every model role at an EU route

.env.local
SUMMARY_MODEL is read by the worker and EMBEDDINGS_MODEL by both the app and the worker, so set them in every process. VECTOR_SIZE defaults to 3584, which matches bge-multilingual-gemma2.
Embeddings already default to Scaleway, but the chat, rephrase and scoring defaults in .env.example are Google Gemini models served through Vertex AI. Leaving REPHRASE_MODEL at its default is the most common miss: it sends every question and the conversation history to Google on every turn.
If you pick gpt-oss-120b for chat, keep MULTIMODAL_FALLBACK_MODEL=mistral-small-3.2 (the default): gpt-oss-120b reads no images, and messages with image or document content are swapped to the fallback.Remove credentials for providers you do not intend to use (VERTEX_*, AWS_*, AZURE_OPENAI_*, OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY, OPENROUTER_API_KEY and so on), so no route can reach them by mistake. Then check with a real call per model:
4

Keep parsing local

.env.local
Without DOCLING_STRICT=1, a Docling failure falls back to the legacy loaders, and the legacy PDF path sends the document to an external model (PDF_MODEL, claude-haiku-4-5 by default).
5

Keep reranking in the EU

.env.local
scaleway is also the provider used when RERANK_PROVIDER is unset. RERANK_PROVIDER=cohere runs Cohere Rerank v3.5 on AWS Bedrock, which is outside EU jurisdiction even in an EU region, unless you point RERANK_COHERE_BASE_URL at a reranker on your own servers.
6

Close the remaining outbound paths

Work through What Still Reaches Outward. For an EU-only setup the ones that matter most:
  • Content moderation calls the OpenAI Moderation API directly and cannot be rerouted. Keep the content-moderation rule off in Guardrails.
  • MCP connectors (Slack, HubSpot, Google and similar) are outbound by design. Enable only the ones you accept.
  • Speech is off by default. SPEECH_PROVIDER=elevenlabs sends audio to ElevenLabs, and SPEECH_PROVIDER=openai sends it to api.openai.com unless you point SPEECH_BASE_URL at your own server.
  • Optional integrations such as Firecrawl (FIRECRAWL_API_KEY) and HeyGen (HEYGEN_API_TOKEN) stay off while their keys are unset. Leave them unset.

Checklist

  • DEFAULT_MODEL, REPHRASE_MODEL, SUMMARY_MODEL, SCORING_MODEL and EMBEDDINGS_MODEL route to EU upstreams, in the app and the worker
  • No credentials for non-EU providers are set
  • npm run gateway:preflight -- --probe passes
  • DOCLING_STRICT=1 is set
  • Reranking is off, uses RERANK_PROVIDER=scaleway, or points at your own server
  • Object storage is RustFS, MinIO / Ceph or an EU provider
  • The content-moderation rule is off; MCP connectors and speech are reviewed
  • Mail, OpenTelemetry and Langfuse point only at infrastructure you run

Trade-offs

We would rather you hear this from us than find it in production.
  • Answer quality. EU-served and open models are good and improving fast, but on some tasks they still answer less well than the strongest commercial models. How much depends on your documents and your questions.
  • Reranking. The Scaleway reranker, qwen3-embedding-8b, is a bi-encoder, which ranks less precisely than a cross-encoder such as Cohere Rerank.
  • Fully local means GPUs. Serving the models yourself removes every external call, but needs GPU capacity. See Local Models.
Measure on your own material before deciding. The eval datasets in apps/web/evals exist for exactly that.

Our Live Demo

demo.ragen.ai answers, searches and reranks with Scaleway Generative APIs: mistral-small-3.2 for answers, bge-multilingual-gemma2 for embeddings and qwen3-embedding-8b for reranking. Uploaded files are kept in Scaleway Object Storage, conversations are encrypted with a key in Scaleway Key Manager, and documents are parsed by Docling with DOCLING_STRICT=1.

Need Help?

Web Amigos built Ragen and deploys it for companies across Europe. If you want an EU-only deployment set up, reviewed or maintained, talk to the team that built it.