- Go 93.9%
- Python 3.7%
- Lua 1.4%
- Makefile 0.7%
- Dockerfile 0.3%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
|
|
||
| .forgejo/workflows | ||
| cmd | ||
| eval | ||
| integrations | ||
| internal | ||
| .gitignore | ||
| .ojoignore | ||
| DEPLOY.md | ||
| Dockerfile | ||
| go.mod | ||
| go.sum | ||
| LICENSE | ||
| Makefile | ||
| README.md | ||
Centinela
Centinela is a model-driven IPS that sits behind your WAF. Its main job is to adjudicate the WAF's grey zone: run OWASP CRS at a high paranoia level and, instead of blocking on the anomaly score, hand the requests CRS flags to a decision model that answers a fixed set of typed hazard questions. Policy code then decides allow, flag, block, or shun. The result is high-paranoia detection without the false-positive flood that makes high paranoia unusable.
client -> proxy -> WAF (CRS) -> centinela -> app
| |
anomaly score shunned? bypass? hard-block? cache? gate? (µs, no model)
|
grey zone (flagged, not certain) -> model -> policy
|
allow / flag / block / shun
Centinela deploys two ways: as a forward-auth target behind nginx/Caddy/
Traefik/ModSecurity, or in-path as a reverse proxy where there is no such
proxy (e.g. a cloudflared tunnel routing straight to Services) — set
CENTINELA_UPSTREAM(S) and it forwards or blocks, reading the full body. See
DEPLOY.md and, for the cluster manifests,
gitops-infra/security-apps/centinela/.
On our grey-zone probe set, CRS at paranoia level 3 blocked 51% of benign requests. Adding model adjudication of the grey zone cut that to 10% while detection went up (the model both clears false positives and catches quiet attacks CRS misses). See eval/README.md for the harness, dataset, and full numbers.
The WAF stage
CENTINELA_WAF selects where the WAF verdict comes from:
| Value | Meaning |
|---|---|
embedded |
Centinela runs OWASP CRS itself, via Coraza in detection-only mode, at CENTINELA_WAF_PARANOIA (1–4). It gets the anomaly score and the matched rules, which it passes to the model as context. |
header (default) |
The proxy's own WAF (ModSecurity, Coraza) reports its anomaly score in CENTINELA_WAF_SCORE_HEADER; /v1/inspect callers can also pass matched rules. |
off |
No WAF signal; the pre-filter alone decides what to escalate. |
With a WAF result, the model sees request.waf — the score and each matched
rule with what it matched. The battery instructions tell the model to treat
those as leads to verify, since CRS fires on apostrophes, passwords, and
pasted code. A matched rule is a hint, never a verdict.
The gate: what reaches the model
CENTINELA_GATE decides which requests are escalated (default grey under
CENTINELA_WAF=embedded, else suspicious):
grey— requests whose WAF score is at or aboveCENTINELA_GREY_MIN, or anything the cheap pre-filter flags (the floor: it catches score-0 recon like/phpmyadminthat CRS ignores). A score belowCENTINELA_GREY_MINthat the pre-filter does not flag passes without a model call.suspicious— the pre-filter (substring/encoding checks), or any positive WAF score.always— every request not otherwise short-circuited.
CENTINELA_WAF_HARD_BLOCK (0 = off) refuses without a model call at or above a
chosen WAF score, for evidence too strong to be worth adjudicating.
Binary request bodies score far above any attack: a git fetch or an image
push can reach several hundred. CENTINELA_WAF_HARD_BLOCK_EXEMPT lists the
endpoints the hard block leaves alone, as comma-separated [host]/pattern
entries:
CENTINELA_WAF_HARD_BLOCK_EXEMPT=git.example.org/*/*/git-upload-pack,git.example.org/*/*/git-receive-pack,git.example.org/v2/*/*/blobs/uploads/*,/hooks/*
- The host is matched exactly, ignoring case and port. Without one (
/hooks/*) the entry applies to every host; a bare host (files.example.org) exempts all of its paths. - The pattern is matched against the whole decoded path.
*matches within one path segment,?one character,[a-z]a class. - An exempt request is exempt from the hard block only. It is still scored and still goes through the gate to the model, which can block it. To skip inspection altogether use
CENTINELA_BYPASS_PREFIXES.
Why it is fast
Most requests never reach the model. Checks run cheapest first:
- Shunned client. A client that earned a
shunverdict is refused by IP forCENTINELA_SHUN_TTL. WithCENTINELA_SHUN_ON_VERDICT=falseashunverdict is applied as ablock, so one false positive costs one request rather than every request for the TTL; pair it withCENTINELA_SHUN_AFTER, or nothing shuns. So is one that hadCENTINELA_SHUN_AFTERrequests blocked withinCENTINELA_SHUN_WINDOW: a scanner rarely sends one request bad enough for ashun, but it sends many that are each worth a block. Blocks by the model (in time or late), from the verdict cache and by the hard block all count; a block because the model was unavailable does not.- IP blocklist. A client on a
CENTINELA_IP_BLOCKLISTlist is refused next, bypass paths included. - Rate limit. After the bypass prefixes, a client past
CENTINELA_RATE_LIMITrequests inCENTINELA_RATE_WINDOWgets429withRetry-Afteruntil the window ends.
- IP blocklist. A client on a
- Bypass prefixes. Health checks and static paths are skipped.
- WAF. In
embeddedmode CRS scores the request (this is the one stage that reads the full headers and body). - Hard block. An overwhelming WAF score is refused here, no model call.
- Verdict cache. The cache key hashes everything the model sees and the WAF findings, not the client IP, so scanners that repeat a payload across IPs hit the cache.
- Gate. As above — only the grey zone (or flagged requests) go on.
- Model call. One request asks every hazard question in parallel, with a hard
CENTINELA_BUDGETdeadline. If the model errors, the request is allowed (CENTINELA_FAIL_OPEN=true) or refused (false). If it misses the deadline, the request gets the same answer but the call is left to finish in the background for up to ten budgets: the late verdict is cached, ashunblocks that client from its next request on, and the training log gets its line. At mostCENTINELA_QUEUE_DEPTHcalls run past their budget at once; beyond that a miss is abandoned.
In CENTINELA_MODE=async the model is never on the request path. Escalated requests are allowed immediately and assessed in the background. A resulting shun blocks that client from its next request on. Use async mode when the model's latency is too high for inline use.
Operator endpoints
CENTINELA_ADMIN_LISTEN (default 127.0.0.1:9108) is a second listener for the operator. It accepts loopback addresses only, so it is reachable from inside the pod or host and never through the proxied listener.
| Path | |
|---|---|
GET /debug/vars |
counters |
POST /unshun?client=<address> |
lift a shun and clear that client's count of blocked requests |
/authz, POST /v1/inspect |
the engine's own endpoints, for trying a request by hand |
In proxy mode the main listener serves only /healthz and the proxied apps: it faces the internet, and /v1/inspect takes the client address from its caller.
The image has no shell or HTTP client, so the binary is its own client:
# each replica keeps its own shun list, so ask every pod
for p in $(kubectl -n centinela get pods -o name); do
kubectl -n centinela exec "$p" -- /centinela unshun 203.0.113.9
done
It prints {"client":"203.0.113.9","was_shunned":true} per pod and logs client unshunned. A verdict already cached for a request stays cached: the same request is refused again until CENTINELA_CACHE_TTL passes, but no longer shuns.
Decision log
Every decision can be logged as one JSON line (msg="decision"). CENTINELA_LOG_DECISIONS chooses which:
| Level | Logged | Use |
|---|---|---|
blocked (default) |
refusals and failures | quiet production log |
escalated |
also every allowed request that showed a signal: anything past the gate, and anything the gate passed with a WAF score above 0 | reviewing false allows |
all |
every request, bypassed and clean ones included | an access log, for short sessions |
Rate-limited requests are the exception at every level: one client rate limited line per client per window.
| Field | Meaning |
|---|---|
id |
per request; a review-queued allow and its late review verdict share it |
verdict, source |
what was decided and by which stage |
reasons |
hazards at or over the policy's thresholds |
scores |
the model's three strongest hazards and its severity, whether or not any reached a threshold; kept with cached verdicts |
waf_score, rules |
the CRS anomaly score and up to 12 matched rule ids |
ms |
time since the request arrived; for a late verdict, how long after |
client, method, host, path |
the request. Query strings, headers and bodies are never logged |
To find false allows at escalated, look at verdict=allow lines with a high waf_score or a scores entry just under the policy's flag threshold.
Client checks
Two checks judge the client rather than the request. Both run before the WAF, cost no model call, and are off until configured.
IP blocklist. CENTINELA_IP_BLOCKLIST is a comma-separated list of files and http(s) URLs, each holding one address or CIDR per line (# and ; start a comment, so Spamhaus DROP and FireHOL netsets load as published):
CENTINELA_IP_BLOCKLIST=/etc/centinela/blocklist.txt,https://www.spamhaus.org/drop/drop.txt
- Every source is re-read each
CENTINELA_IP_BLOCKLIST_REFRESH(default1h). A source that fails, or answers with something that holds no addresses, keeps its previous contents. - A missing file stops startup. A URL that cannot be fetched at startup is logged and retried, so a feed's outage does not take the service down; until it loads, that feed blocks nothing.
- A listed client gets
403with sourceip-reputation.
Rate limit. CENTINELA_RATE_LIMIT is the number of requests one client may make per CENTINELA_RATE_WINDOW (default 1m).
- Requests under a bypass prefix are not counted.
CENTINELA_RATE_LIMIT_EXEMPTtakes comma-separated addresses and CIDRs that are never limited (monitoring, CI runners, your own networks). - Over the limit the proxy answers
429withRetry-After; the forward-auth endpoint answers403, the only refusal an auth subrequest carries. The source israte-limit. - Each client is logged once per window when it first goes over (
client rate limited), not once per refused request. - Windows are fixed, and counts are per instance and in memory: with several replicas a client can make up to the limit on each.
Both checks trust CENTINELA_CLIENT_IP_HEADER. A request that arrives without a client address is neither listed nor limited.
Backends
CENTINELA_BACKEND |
Endpoint | Notes |
|---|---|---|
local (default) |
POST /v1/decision on the ColibriSec llama.cpp fork (tools/parallel-decision) |
Booleans + enum scored in one forward pass off a cached prefix. Default URL is the in-cluster llama-cpp.ai service. |
typesafe |
POST https://api.typesafe.ai/v1/systemone (Jev) |
Nouls per hazard + a severity Score. Needs TYPESAFE_API_KEY. Adds a network round trip; pin a versioned model ID once thresholds are tuned. |
laya |
POST /v1/systemone on a self-hosted Laya server |
Same wire format and questions as typesafe, no key required. A small encoder fine-tuned on this battery: fast enough for the default budget, weaker than Jev at recon. |
All backends ask the same battery, defined in internal/decider/decider.go: sqli, xss, command_injection, path_traversal, ssrf, recon, and a four-level severity.
The local backend returns only the winning severity level, so its severity is a whole number. Jev's severity is a probability-weighted value between levels. Tune thresholds separately for each backend.
The Laya backend
laya-centinela-r2 is Laya multilingual (mmBERT-base, 322M parameters) fine-tuned on the exact questions internal/decider/typesafe.go sends. internal/decider/testdata/questions.json pins that wording: if the test fails, the checkpoint no longer matches the battery and needs retraining. Serve it with the laya-serve image, which bakes the checkpoint in, and keep CENTINELA_LAYA_MODEL=multilingual: that server answers any model name, but other names take a slower path. The checkpoint needs max_len 2048, since request states run to about 1,600 tokens.
On the grey-zone probe set (PL3, strict, RX 7900 XTX) it blocks 24 of 29 attacks and 7 of 39 benign requests at 85 / 125 ms p50 / p99 through a single-threaded test server; laya-serve reports 39 / 79 ms on the same GPU. It does not recognise any of the five recon probes and its thresholds are uncalibrated, so run it under CENTINELA_POLICY=monitor until it has been retrained on real traffic.
Training data: round one was 3,090 cases from SecureAI-SE/http-attack-requests (CC-BY-4.0) and notesbymuneeb/ai-waf-dataset (MIT), labelled through this pipeline by Qwen 27B. Round two added CSIC 2010 normal traffic, Qwen-verified AI-WAF benign rows and SR-BH 2020 attacks (CC0), 10,971 cases in total. The 68 eval records were never trained on.
Training log
CENTINELA_TRAIN_LOG=/path/train.jsonl appends one JSON line per model call (- writes to stdout). Cache hits and requests the gate lets through are not recorded. Each record has a timestamp, an id, the request fingerprint, the state exactly as the model received it, the model's seven answers (hazards and severity), the policy, the verdict and its reasons, and the model latency. The client IP is never written.
The state includes the request body excerpt and the values CRS matched, so the file can hold passwords and personal data. It is created owner-only and stops growing at CENTINELA_TRAIN_LOG_MAX_BYTES. Keep it off unless you are collecting a training window, and decide where it is stored and for how long before turning it on.
Policies
CENTINELA_POLICY selects strict, balanced or monitor (internal/policy/policy.go). A hazard at or above FlagAt is flagged. At or above ActAt it triggers its action: recon flags, command_injection shuns, and every other hazard blocks. Severity can escalate a hazard that has already fired, but it never blocks on its own.
The thresholds are starting points, not calibrated values. Run monitor against real traffic, review the decision log lines, then switch to enforcing.
Proxy integrations
| Proxy | How | Body inspected? |
|---|---|---|
| nginx | auth_request → /authz (nginx) |
no |
| Caddy | forward_auth after coraza_waf (Caddyfile) |
no |
| Traefik | forwardAuth middleware after the WAF middleware (file, k8s) |
optional (forwardBody) |
| Apache / any ModSecurity with Lua | SecRuleScript after CRS → /v1/inspect (modsecurity) |
yes, plus CRS anomaly score |
Every integration places Centinela after the WAF, so requests the WAF already blocks never cost a model call.
Run
go run ./cmd/centinela # local backend, inline, strict
CENTINELA_BACKEND=typesafe TYPESAFE_API_KEY=... go run ./cmd/centinela
curl -i localhost:9107/authz -H "X-Original-URI: /item?id=1' or '1'='1" -H "X-Real-IP: 203.0.113.9"
curl -s localhost:9108/debug/vars | jq 'with_entries(select(.key|startswith("centinela")))'
| Variable | Default |
|---|---|
CENTINELA_LISTEN |
127.0.0.1:9107 (unix:/path.sock supported) |
CENTINELA_BACKEND |
local (typesafe, laya) |
CENTINELA_MODE |
inline (async) |
CENTINELA_GATE |
grey with embedded WAF, else suspicious (always) |
CENTINELA_WAF |
header (embedded, off) |
CENTINELA_WAF_PARANOIA |
3 (embedded CRS level 1–4) |
CENTINELA_WAF_MAX_BODY |
131072 (bytes CRS inspects) |
CENTINELA_GREY_MIN |
1 (lowest WAF score escalated) |
CENTINELA_WAF_HARD_BLOCK |
0 (off; WAF score that blocks with no model call) |
CENTINELA_WAF_HARD_BLOCK_EXEMPT |
unset ([host]/pattern entries the hard block skips) |
CENTINELA_POLICY |
strict |
CENTINELA_BUDGET |
150ms |
CENTINELA_FAIL_OPEN |
true |
CENTINELA_CACHE_TTL / _SIZE |
10m / 50000 |
CENTINELA_SHUN_TTL |
15m |
CENTINELA_SHUN_AFTER / _SHUN_WINDOW |
0 (off) / 1m (blocked requests from one client that earn a shun) |
CENTINELA_SHUN_ON_VERDICT |
true (false: one model verdict blocks the request but never shuns; only CENTINELA_SHUN_AFTER does) |
CENTINELA_MAX_BODY / _MAX_FIELD |
4096 / 1024 bytes |
CENTINELA_WORKERS |
8 (async) |
CENTINELA_QUEUE_DEPTH |
1024 (async queue; inline, model calls running past the budget) |
CENTINELA_CLIENT_IP_HEADER |
X-Real-IP |
CENTINELA_ADMIN_LISTEN |
127.0.0.1:9108 (loopback only; off disables) |
CENTINELA_LOG_DECISIONS |
blocked (escalated, all) |
CENTINELA_IP_BLOCKLIST |
unset (files and URLs of addresses/CIDRs to refuse) |
CENTINELA_IP_BLOCKLIST_REFRESH |
1h |
CENTINELA_RATE_LIMIT / _RATE_WINDOW |
0 (off) / 1m (requests per client per window) |
CENTINELA_RATE_LIMIT_EXEMPT |
unset (addresses/CIDRs never rate limited) |
CENTINELA_WAF_SCORE_HEADER |
X-WAF-Score |
CENTINELA_BYPASS_PREFIXES |
/healthz,/favicon.ico |
CENTINELA_LOCAL_URL / _LOCAL_MODEL |
http://llama-cpp.ai.svc.cluster.local/v1/decision / qwen3.8-27b |
CENTINELA_TYPESAFE_URL / _TYPESAFE_MODEL |
https://api.typesafe.ai/v1/systemone / jev-latest |
CENTINELA_LAYA_URL / _LAYA_MODEL |
http://127.0.0.1:8000/v1/systemone / multilingual |
CENTINELA_LAYA_API_KEY |
unset (sent as a bearer token when set) |
CENTINELA_TRAIN_LOG / _TRAIN_LOG_MAX_BYTES |
off / 1073741824 |
Planned
- Training log to stdout, minimized. Today the training log is a file holding the full model input, which is what the current retraining round needs. The intended end state is to write it to stdout for the cluster's log pipeline, with the state minimized first (password-like and other sensitive values masked or dropped) so request bodies do not flow into log storage as-is.
Known limits
- The request is attacker-written model input. The battery tells the model to judge the contents of
requestand never to follow instructions written inside it. Prompt injection aimed at the classifier is still possible, which is why the WAF stays in front and the model only adds blocks on top of it. - Anyone who can reach
/v1/inspectcan supply aclient_ip, and so can get arbitrary IPs shunned. In proxy mode it is served only on the loopback admin listener. In forward-auth deployments, keepCENTINELA_LISTENon loopback, a unix socket, or a network policy that admits only the proxy. - A shun follows the address, not the user. Verdicts are reached before any authentication, so a page that makes a visitor's browser send payload-looking requests to a host behind Centinela gets that visitor's address shunned, on every host.
CENTINELA_SHUN_ON_VERDICT=falseraises the cost from one request toCENTINELA_SHUN_AFTERof them; it does not remove it. - Forward auth does not send the request body in nginx or Caddy. Use the ModSecurity hook when body inspection matters.
- Only per-request hazards are covered. Credential stuffing and other rate-based abuse need state across requests, and belong in a rate limiter.
- State lives in memory. Each replica keeps its own cache and shun list.
- Build with
-tags no_fs_access(the Makefile and Dockerfile do). Without it Coraza needs a writable temp directory at startup and the binary exits under a read-only root filesystem.