LLM gateway
The gateway is not enabled on this server. It is part of the business plan and of every self-hosted installation.
The gateway sits between your application and the AI provider. It behaves exactly like the OpenAI API, so you only change the base URL. For every request it replaces names, addresses, ID numbers and other personal data with realistic fake values, sends the request on, and puts the real values back into the answer. The provider never sees the real data.
- Your app sends
“Write to Janez Novak …” - The provider receives
“Write to Klemen Mlakar …” - The provider answers
“Dear Mr Mlakar …” - Your app receives
“Dear Mr Novak …”
Quick start
Use any OpenAI SDK. Base URL: http://pii.kwizmo.eu/gateway.php/v1 (once enabled). On Apache the shorter http://pii.kwizmo.eu/v1 works too.
from openai import OpenAI
client = OpenAI(
base_url="http://pii.kwizmo.eu/gateway.php/v1", # the only change
api_key="pii_…", # your Kwizmo business key
)
reply = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content":
"Write a short reply to Janez Novak (janez.novak@gmail.com) about his contract."}],
)
print(reply.choices[0].message.content) # contains the real name and address againKeys
Every request needs a Kwizmo API key with the gateway permission (see business plans). There are two ways to pay for the AI model:
- Provider key stored by the operator: send only your Kwizmo key,
Authorization: Bearer pii_…. - Your own provider key: send it as usual in
Authorizationand the Kwizmo key inX-PII-Key. It is passed on and never stored.
client = OpenAI(
base_url="http://pii.kwizmo.eu/gateway.php/v1",
api_key="sk-…", # your own provider key, passed on unchanged
default_headers={"X-PII-Key": "pii_…"}, # your Kwizmo key
)Supported endpoints
| Endpoint | What happens |
|---|---|
POST /v1/chat/completions | All messages are protected (text, content parts, tool calls and tool results). The answer is restored, including streamed answers and tool call arguments. |
POST /v1/embeddings | The input texts are protected before they are embedded. |
GET /v1/models | Models of all configured providers. |
Images and files inside messages are passed on unchanged: the gateway cannot read text in pictures.
Options
Send options as HTTP headers, or in a "pii" object in the request body (the object is removed before forwarding). All are optional.
| Header | pii field | Meaning |
|---|---|---|
X-PII-Locale | locale | Language of the texts: auto, en, de, at, ch, sl, hr, sr, bs, me, da, no, sv, fi, is. Default: auto. |
X-PII-Disable | disable | Detectors to skip, comma-separated in the header: name, ml_name, address, postcode, phone, email, iban, national_id, credit_card, date_of_birth, account, secret, ip, ssn. |
X-PII-Known-Names | known_names | JSON array of names to always replace, e.g. your customers. |
X-PII-Consistency-Key | key | Same key → same fake values in every request, useful for long conversations sent in parts. |
X-PII-Restore | restore | false returns the answer with the fake values. |
X-PII-Return-Mapping | return_mapping | true adds "pii": {"replacements", "mapping"} to non-streamed answers. |
X-PII-Instruction | instruction | false skips the short system message that asks the model to copy names and numbers exactly. |
Every answer has the header X-PII-Replacements with the number of replaced values.
More examples
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://pii.kwizmo.eu/gateway.php/v1", apiKey: process.env.KWIZMO_KEY });
const stream = await client.chat.completions.create({
model: "gpt-4o-mini",
stream: true,
messages: [{ role: "user", content: "Povzemi pritožbo gospe Maje Kovačič, tel. 041 123 456." }],
pii: { locale: "sl", known_names: ["Maja Kovačič"] }, // optional, removed before forwarding
});
for await (const chunk of stream) process.stdout.write(chunk.choices[0]?.delta?.content ?? "");curl -s http://pii.kwizmo.eu/gateway.php/v1/chat/completions \
-H "Authorization: Bearer pii_…" \
-H "Content-Type: application/json" \
-H "X-PII-Locale: hr" \
-H "X-PII-Return-Mapping: true" \
-d '{"model": "gpt-4o-mini",
"messages": [{"role": "user", "content": "Pošalji podsjetnik Ivani Marić, OIB 69435151530."}]}'<?php
// PHP 7.4+, no libraries needed
$ch = curl_init('http://pii.kwizmo.eu/gateway.php/v1/chat/completions');
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => ['Content-Type: application/json', 'Authorization: Bearer ' . getenv('KWIZMO_KEY')],
CURLOPT_POSTFIELDS => json_encode([
'model' => 'gpt-4o-mini',
'messages' => [['role' => 'user', 'content' => 'Sehr geehrte Frau Müller, …']],
]),
]);
$reply = json_decode(curl_exec($ch), true);
echo $reply['choices'][0]['message']['content'];Providers
The gateway works with any service that offers the OpenAI chat API, for example OpenAI, Azure OpenAI, Mistral (EU), Anthropic (OpenAI-compatible endpoint), Google Gemini (OpenAI-compatible endpoint), Groq, OpenRouter, and local models with Ollama, vLLM or LM Studio. On a self-hosted installation the operator chooses providers per model in app/conf.php:
'gateway_enabled' => true,
'gateway_upstreams' => [
// first match wins; "models" are prefixes, "*" matches everything
['name' => 'mistral', 'base_url' => 'https://api.mistral.ai/v1', 'api_key' => 'MISTRAL_KEY', 'models' => ['mistral-', 'codestral']],
['name' => 'local', 'base_url' => 'http://127.0.0.1:11434/v1', 'no_key' => true, 'models' => ['llama', 'qwen']],
['name' => 'openai', 'base_url' => 'https://api.openai.com/v1', 'api_key' => 'OPENAI_KEY', 'models' => ['*']],
],What is replaced and restored
- The same person gets the same fake name throughout the conversation, also in other grammatical cases (“Novak”, “Novaka”, “Novakom”).
- When the model uses a fake name in a form that was not in your text (“Mlakarju”), the real name is still restored in the matching form (“Novaku”).
- Streamed answers are held back only for as long as a fake value could still be incomplete, usually a few words.
- If the model changes a fake value (translates, abbreviates or misspells it), that spot stays fake. The system message reduces this; check important answers.
- Detection is automatic and not perfect: see the measured accuracy.
Errors and limits
Errors use the OpenAI format, so SDKs raise their usual exceptions. Errors from the provider are passed on unchanged.
| Status | Code | Meaning |
|---|---|---|
| 401 | api_key_required, invalid_api_key | Missing or wrong Kwizmo key. |
| 401 | provider_key_required | No provider key stored for this model; send your own. |
| 403 | api_key_scope | The key is not allowed to use the gateway. |
| 400 | model_not_available | No provider is configured for this model. |
| 413 | payload_too_large, text_too_long | Request larger than your key allows. |
| 429 | rate_limited, quota_exceeded | Per-minute limit or monthly quota of the key reached. See Retry-After, X-RateLimit-* and X-Quota-*. |
| 502 | upstream_unreachable | The provider did not answer. |
Privacy
Requests and answers are processed in memory only. The gateway writes nothing about their content; it logs the time, the IP address and the API key id, and counts requests per key for billing. The provider receives only the protected text and remains a separate recipient under its own terms. For contracts with business customers we sign a data processing agreement.
