Behavioral Baselines
Behavioral baselines help compare an agent’s actions with expected task behavior. They can support detection and regression analysis, but a deviation is not proof of compromise and an apparently normal run can still be harmful.
Detection & Monitoring · security-teams · builders Why Baselines
A hijacked agent may use the same tools and interfaces as a benign run. Compare task, identity, data access and destination as well as sequence and volume. Policy violations can sometimes be detected directly, even without a learned behavioral baseline.
Collect the events needed to investigate actions: task, caller, tool, resource, destination, policy decision and outcome. Redact sensitive content, control access and set retention limits. Internal chain-of-thought may be unavailable and is not a prerequisite for useful auditability.
What to Monitor
Four layers of behavioral signals, from easiest to hardest to implement.
Which tools the agent calls, in what order, how often, and with what parameters. A code review agent that suddenly calls web_fetch or reads .ssh/ has deviated from its baseline. Tool events are useful observable signals. A deviation can be legitimate; an unauthorized action can also resemble the baseline.
Sudden spikes in token consumption, unusual response times, or cost anomalies. Some harmful runs consume extra resources; others use a normal or smaller number of calls. Cost spikes can also have benign causes. This is cheap to monitor and catches resource exhaustion attacks (Kill Chain Stage 4).
What the agent outputs — does it contain data the task didn't request? PII, credentials, file contents that weren't part of the assignment? Output classifiers can detect when agent responses contain unexpected sensitive data — catching Stage 5 EXFILTRATE at the output boundary.
Model scores can inform triage only after calibration on the relevant task distribution. The previous 90%/60% execution thresholds were heuristics, not validated probabilities. Risk, permission and irreversible impact should determine whether an action requires human approval even when a score is high.
Record the matched rule, evidence and review status separately from any model score. Deterministic means the same input produces the same rule result; it does not mean 100% accuracy. A field named confidenceScore is not a calibrated probability unless evaluation supports that interpretation.
Five Tools
The observability stack for agent behavioral monitoring. Includes open standards, open-source tools and vendor platforms. Instrumentation is not a security detector by itself.
OpenTelemetry publishes evolving semantic conventions for generative AI and agent telemetry. Pin the version used by your instrumentation and backend; event names and stability levels can change. Record supported agent/tool spans, latency and token usage without indiscriminate capture of sensitive content. Current conventions.
LangSmith records application traces and supports evaluations and feedback. Configure which inputs and outputs are captured and how they are protected. Evaluate the actual instrumented tools; coverage and security detection do not follow automatically from using an observability platform.
Open-source LLM observability built on OpenTelemetry. Traces agent runs, tool calls, and model request/response with full context. Supports evaluation via LLM-based evaluators, code-based checks, or human labels. Integrates with Claude Agent SDK, OpenAI Agents SDK, LangGraph, and CrewAI.
Driftbase compares behavioral fingerprints across runs or versions, including tool paths and latency. Its documentation recommends a minimum sample for a verdict. A drift score identifies change, not whether the change is malicious or operationally unacceptable.
Complete observability for the AI stack — LLMs, vector databases, GPUs. One line of code to instrument. Built on OpenTelemetry, so traces are compatible with any OTel-compliant backend.
The 3-Pass Pattern
This proposed three-pass workflow combines per-file review, cross-file analysis and independent challenge. Each pass has a different purpose. Assess differences against evidence; they do not by themselves establish compromise.
Scan individual files and attach paths, locations, rule identifiers and evidence to findings. Zero findings can be a valid result. Measure coverage against known issues and clean examples rather than expecting a fixed finding density.
Analyze interactions across files and trust boundaries. This pass should be able to discover new issues that no single-file review reveals. Check the evidence and explain why the issue requires the wider context.
Independently challenge earlier findings and search for missed cases. New findings or disagreements can reflect better coverage, ambiguous requirements or errors. Investigate those possibilities before attributing differences to injected context.
Honest Limitations
Agent behavior changes legitimately over time — new tools added, workflows updated, models upgraded. Your baseline becomes stale. You need to re-calibrate regularly, which means you need to distinguish "the agent evolved" from "the agent was compromised." This is the hardest problem in behavioral monitoring.
The sources reviewed here do not establish comparable production false-positive and false-negative rates across these tools. Research such as SentinelAgent explores behavioral oversight, but does not validate this proposed workflow. Ask for datasets, labels, task coverage, thresholds and uncertainty before relying on detection claims.
Small changes may remain below a particular alert threshold or contaminate an automatically updated baseline. Test slow attacks, preserve a trusted reference set, and combine behavioral comparison with explicit permission and destination policies.
Behavioral baselines are the detection layer in the Kill Chain. Combine with hook-based guardrails (prevention), MCP security (tool defense), and red teaming (validation).
Governance guides, more detection patterns, and practitioner content coming.
This work represents the author's independent research and personal views. It is not related to or endorsed by the author's employer.