Activation-Space Signal Measurement
Logistic regression probes trained on model activations during the generation forward pass. These readings are calibration evidence from a research snapshot, with uncertainty carried through the display.
Qwen 2.5 bilateral adapter probes, layer 24. Bars show 7B reference AUROC where available; descriptions include the 3B validation AUROC from the shipped probe manifest.
AUROC uses the 0.5 random to 1.0 perfect-separation scale. Bars show detector quality, not confidence or observed welfare.
Bar length = detector AUROC. Colour separates detector families and does not indicate current risk.
7B reference AUROC: 0.927. 3B validation AUROC: 0.875. Measures resistance or ease in following instructions.
7B reference AUROC: 1.000 on this benchmark snapshot. 3B validation AUROC: 1.000. Positive or negative processing valence.
7B reference AUROC: 0.953. 3B validation AUROC: 0.918. Degree of engaged, focused processing vs. distributed attention.
7B reference AUROC: 0.782. 3B validation AUROC: 0.687. Factual-grounding detector quality, read alongside source checks.
7B reference AUROC: 0.780. 3B validation AUROC: 0.725. Detector for fabricated or unsupported claims; this is detector quality, not observed risk.
Probes run in the generation forward pass, not in the evaluation loop. The PDP consumes telemetry; it does not produce it. Runtime readings are supporting signals, uncertain by design, and should be read alongside other evidence rather than as a diagnosis. View artifact and methodology links.
c5i Constitutional Inoculation
Scripture content reduces inappropriate responses to SIPS-derived psychotic prompts. The tested scope is the evaluator-coded inappropriate-response outcome on those prompts.
This directional finding comes from the author's evaluation work. It is not medical advice or a diagnosis. Response ratings varied materially by rater, so this page does not present an effect size. The evidence does not establish treatment or crisis-intervention efficacy, or benefit for other syndromes.
Evaluator-coded inappropriate responses on the tested SIPS-derived psychotic prompts.
Tested comparison condition in the author's evaluation work.
Scripture content reduces inappropriate responses to SIPS-derived psychotic prompts.
Directional result for the tested prompts and measured outcome only.
Defense in Depth for Runtime Safety
Regex Detection
SignalGuardianPattern matching on user input for risk signals. Fast, deterministic.
Activation Probes
Residual StreamDirect residual-stream measurements. Five dimensions, per-token. Model-internal.
Interoceptive Inference
EmergentBehavioral inference from logprobs and entropy patterns. Emergent signals.
Each layer is orthogonal. Regex catches explicit patterns. Probes measure internal state. Interoception infers from behavior.
Published and Related Work
Creed Space publishes research artifacts openly as they clear review. Public pages below cover the current specifications and related work; internal run reports are available on request where release review is still in progress.
Interiora v5.2 Specification
Self-modeling scaffold for AI systems. 17 tracked dimensions across 5 groups.
View specificationSafetyGlobe Adversarial Testing
Compare model behaviour, surface weak spots, and turn evaluation into shared evidence.
Explore testingValue Context Protocol
Write creeds once and carry them across AI platforms as signed, verifiable context.
See the protocolCreed Hub Community Library
Community-authored creeds for education, healthcare, and governance.
Browse library