Add Laya backend and an opt-in training log #1

Merged
james merged 5 commits from laya-backend into main 2026-10-02 21:18:46 +00:00
Owner

Adds a laya decision backend and an opt-in training log, so Centinela can run the fine-tuned Laya model in monitor mode and collect real traffic for a third training round.

What changes

  • CENTINELA_BACKEND=laya: calls a self-hosted Laya server at CENTINELA_LAYA_URL. Laya speaks the Jev /v1/systemone wire format, so this reuses the TypeSafe decider; the API key is optional (CENTINELA_LAYA_API_KEY).
  • Question wording is pinned: internal/decider/testdata/questions.json is the exact set the checkpoint was trained on, and a test fails if the battery drifts from it.
  • CENTINELA_TRAIN_LOG: one JSON line per model call with the exact state sent, the seven answers, verdict, reasons, id, fingerprint, timestamp and latency. Off by default, owner-only file, stops at CENTINELA_TRAIN_LOG_MAX_BYTES (1 GiB). The client IP is never written. Cache hits and gated requests are not recorded.
  • centinela-eval -backend laya, plus README / DEPLOY updates.

Privacy

The logged state is the full model input, so it includes the body excerpt and CRS-matched values and can hold passwords and personal data. That is intended (password-like values are one of the model's main false-positive classes); the file should live on a dedicated volume and be removed when the collection window ends.

Verification

  • make vet, make test (race) and make cover pass locally; total coverage 38.5% -> 44.3%.
  • centinela-eval -backend laya against laya-centinela-r2 on the RX 7900 XTX: 24/29 attacks blocked, 7/39 benign blocked, 85 / 125 ms p50 / p99.
  • Ran the binary with CENTINELA_BACKEND=laya CENTINELA_WAF=embedded CENTINELA_POLICY=monitor CENTINELA_BUDGET=150ms and the training log on: model calls were recorded, cache hits and gated requests were not, no client IP or cookie in the file.

Not in this PR

  • No cluster manifests. Deploying needs a Laya model service (checkpoint hosting and GPU are undecided) and a volume for the training log.
  • The model misses all five recon probes in the eval and its thresholds are uncalibrated: keep CENTINELA_POLICY=monitor.
Adds a `laya` decision backend and an opt-in training log, so Centinela can run the fine-tuned Laya model in monitor mode and collect real traffic for a third training round. ## What changes - **`CENTINELA_BACKEND=laya`**: calls a self-hosted Laya server at `CENTINELA_LAYA_URL`. Laya speaks the Jev `/v1/systemone` wire format, so this reuses the TypeSafe decider; the API key is optional (`CENTINELA_LAYA_API_KEY`). - **Question wording is pinned**: `internal/decider/testdata/questions.json` is the exact set the checkpoint was trained on, and a test fails if the battery drifts from it. - **`CENTINELA_TRAIN_LOG`**: one JSON line per model call with the exact `state` sent, the seven answers, verdict, reasons, id, fingerprint, timestamp and latency. Off by default, owner-only file, stops at `CENTINELA_TRAIN_LOG_MAX_BYTES` (1 GiB). The client IP is never written. Cache hits and gated requests are not recorded. - `centinela-eval -backend laya`, plus README / DEPLOY updates. ## Privacy The logged `state` is the full model input, so it includes the body excerpt and CRS-matched values and can hold passwords and personal data. That is intended (password-like values are one of the model's main false-positive classes); the file should live on a dedicated volume and be removed when the collection window ends. ## Verification - `make vet`, `make test` (race) and `make cover` pass locally; total coverage 38.5% -> 44.3%. - `centinela-eval -backend laya` against `laya-centinela-r2` on the RX 7900 XTX: 24/29 attacks blocked, 7/39 benign blocked, 85 / 125 ms p50 / p99. - Ran the binary with `CENTINELA_BACKEND=laya CENTINELA_WAF=embedded CENTINELA_POLICY=monitor CENTINELA_BUDGET=150ms` and the training log on: model calls were recorded, cache hits and gated requests were not, no client IP or cookie in the file. ## Not in this PR - No cluster manifests. Deploying needs a Laya model service (checkpoint hosting and GPU are undecided) and a volume for the training log. - The model misses all five recon probes in the eval and its thresholds are uncalibrated: keep `CENTINELA_POLICY=monitor`.
Add Laya backend and an opt-in training log
Some checks failed
test / go (pull_request) Has been cancelled
security-scan / security-scan (pull_request) Has been cancelled
6fb1aff956
Laya is a small encoder fine-tuned on Centinela's seven questions. It
speaks the Jev /v1/systemone wire format, so the backend reuses the
TypeSafe decider with its own URL and an optional API key.

- config: CENTINELA_BACKEND=laya with CENTINELA_LAYA_URL, _LAYA_MODEL
  and optional CENTINELA_LAYA_API_KEY
- decider: testdata/questions.json pins the question wording the
  checkpoint was trained on; the Authorization header is only sent when
  a key is set
- trainlog: CENTINELA_TRAIN_LOG appends one JSON line per model call
  with the exact state sent, the seven answers, verdict, id and
  timestamp, for re-labelling and retraining. Owner-only file, capped
  by CENTINELA_TRAIN_LOG_MAX_BYTES (1 GiB), no client IP. Off by default
- centinela-eval: -backend laya

On the probe set (PL3, strict, RX 7900 XTX) laya-centinela-r2 blocks
24/29 attacks and 7/39 benign at 85 / 125 ms p50 / p99. It misses all
five recon probes, so it is for monitor mode until retrained.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
README: plan to move the training log to stdout with minimization
Some checks failed
security-scan / security-scan (pull_request) Failing after 23s
test / go (pull_request) Successful in 4m22s
83ae7eefb9
Co-Authored-By: Claude Opus 5.5 <[email protected]>

🔎 ojo scan results

Severity Count
🟢 LOW 1
Details (1)
Type Severity ID/Rule Location Description
misconfig 🟢 LOW dockerfile-no-healthcheck Dockerfile:1 image has no HEALTHCHECK
<!-- ojo-scan-summary --> ### 🔎 ojo scan results | Severity | Count | |---|---| | 🟢 LOW | 1 | <details><summary>Details (1)</summary> | Type | Severity | ID/Rule | Location | Description | |---|---|---|---|---| | misconfig | 🟢 LOW | dockerfile-no-healthcheck | <a href="https://git.colibrisec.org/ColibriSec/centinela/src/commit/ae3fa73843e6a65de1b74e8573d897e1d15d7f6b/Dockerfile#L1" target="_blank" rel="noopener noreferrer">Dockerfile:1</a> | image has no HEALTHCHECK | </details>
Dockerfile: set the nonroot user explicitly
All checks were successful
security-scan / security-scan (pull_request) Successful in 1m35s
test / go (pull_request) Successful in 3m10s
c93f7c80f7
The distroless nonroot base already runs as 65532; the scanner flags the
image as root unless USER is stated.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Default the laya backend to laya-serve's port and model name
Some checks failed
security-scan / security-scan (pull_request) Has been cancelled
test / go (pull_request) Has been cancelled
891d883cab
laya-serve answers on port 8000 and is fastest when asked for
"multilingual"; other model names give the same answers on a slower
path.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
gofmt config.go
All checks were successful
security-scan / security-scan (pull_request) Successful in 1m46s
test / go (pull_request) Successful in 2m14s
ae3fa73843
Co-Authored-By: Claude Opus 5.5 <[email protected]>
james merged commit 642250ac69 into main 2026-10-02 21:18:46 +00:00
james deleted branch laya-backend 2026-10-02 21:19:11 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
ColibriSec/centinela!1
No description provided.