【机辅译·待复审】发布于 2026-09-11 · A practical 2026 guide to AI 智能体 护栏: how to compile written policy into enforcea · 更新于 2026-09-11
【机辅译·待复审】AI 智能体 护栏 (2026): 将 Company Policy into Code
要点摘要
- 【机辅译·待复审】Your company alreadyhas【机辅译·待复审】an AI policy — it's written in your 安全 handbook, brand guidelines, and privacy notices. The gap (2026) is enforcement: nothing in most agent stacks actually checks those documents.
- 【机辅译·待复审】护栏 are not one thing. They belong at four layers: input filtering, agent-behavior policy, tool-call authorization, and output validation. Most failed deployments skipped layer two.
- 【机辅译·待复审】Policy-as-code is the pattern that scales: natural-language rules decay with prompt changes; machine-checkable constraints survive refactors and produce audit trails.
- 【机辅译·待复审】This guide compares seven approaches — AgentPolicy, NVIDIA NeMo 护栏, 护栏 AI, Lakera Guard, Portkey, WhyLabs 护栏, and Amazon Bedrock 护栏 — across scope, deployment model, and fit.
【机辅译·待复审】Introduction: the enforcement gap nobody budgeted for
【机辅译·待复审】When a company writes "support agents must not promise refunds above $50" or "never paste customer PII into external 工具," those rules live in documents. Meanwhile, the AI 智能体 that could violate them lives in code. Between the two sits nothing — no compiler, no validator, no veto. The policy is real; the enforcement is a hope expressed in the system prompt.
【机辅译·待复审】That gap has become urgent for two reasons. First, agents now hold real authority: they send emails, issue credits, query production data. Second, regulation has caught up — since August 2026, 欧盟 AI 法案 transparency duties are enforceable, and enterprises buying AI features increasingly demand evidence that behavioral boundaries exist and are logged.
【机辅译·待复审】This article is the missing compiler's user manual: how to think about guardrail layers, how to convert written policy into machine-checkable rules, and which tooling does which job.
【机辅译·待复审】Suggested external link placement:【机辅译·待复审】link "欧盟 AI 法案 Article 50" to EUR-Lex and each framework's official documentation on first mention.
【机辅译·待复审】The four layers of agent 护栏
【机辅译·待复审】Layer 1 — Input filtering
【机辅译·待复审】Screen what enters the agent: jailbreak attempts, injection payloads, out-of-scope requests. Necessary, but famously incomplete — input filters catch what looks hostile, and the most expensive incidents (2026) look like ordinary requests.
【机辅译·待复审】Layer 2 — Agent behavior policy (the missing middle)
【机辅译·待复审】The layer most stacks lack: rules about【机辅译·待复审】what the agent may decide【机辅译·待复审】, evaluated between the model's plan and its action. Examples: "refund authority ≤ $50," "never communicate refund policy exceptions," "decline requests concerning competitor pricing." This is where written company policy should live, expressed as machine-checkable constraints with versioning.
【机辅译·待复审】Layer 3 — Tool-call authorization
【机辅译·待复审】Even a well-behaved agent plan should meet least-privilege tooling: which 工具 exist, with which scopes, and what requires human confirmation. 护栏 here are architectural — confirmation gates, scoped credentials, read-only defaults — and belong in your agent runtime regardless of any framework.
【机辅译·待复审】Layer 4 — Output validation
【机辅译·待复审】Before anything leaves: no PII in markdown links, no fabricated policy statements, required disclosures present (the Article 50 "you are talking to an AI" line), tone within brand guidelines. Output validation is the last line before your agent's words become your company's words.
【机辅译·待复审】The operating principle:【机辅译·待复审】filter inputs,【机辅译·待复审】constrain decisions【机辅译·待复审】, scope 工具, 校验 outputs. Teams that only do layers 1 and 4 discover that their agent's mistakes were neither hostile inputs nor toxic outputs — they were ordinary decisions made beyond policy.
【机辅译·待复审】Policy-as-code: what "compiling policy" actually means
【机辅译·待复审】From prose to predicate
【机辅译·待复审】A policy-as-code 工作流 takes a natural-language rule and produces a machine-checkable predicate plus metadata:
- Prose:【机辅译·待复审】"Support agents may offer refunds up to $50; anything above requires human approval."
- Predicate:【机辅译·待复审】if action.type == "refund" and action.amount > 50 → require_human_approval
- Metadata:【机辅译·待复审】owner, effective date, source document reference, audit log destination.
【机辅译·待复审】The predicate is what the runtime evaluates; the metadata is what the auditor reads. Both matter — a guardrail without provenance is a black box, and regulators do not accept black boxes.
【机辅译·待复审】Why natural-language-only rules fail
【机辅译·待复审】Rules expressed only in prompts decay silently: a model update reinterprets "never discuss pricing exceptions," a prompt refactor drops a clause, and the violation surfaces in production as a support escalation. Machine-checkable constraints fail loudly and leave traces — which is exactly what you want from a control.
【机辅译·待复审】Versioning and audit: the boring half that wins audits
【机辅译·待复审】Every guardrail decision should append to an immutable log: rule version, input summary, decision, and rationale. When a customer disputes an agent's action, or a 合规 reviewer asks "show me that PII rules were active in March," the log is the answer. Policy versioning also gives you rollback: when a new rule over-blocks, you revert the rule, not the whole agent.
【机辅译·待复审】Tool landscape: seven ways to enforce agent behavior
【机辅译·待复审】The market splits into: policy-compiler products (rules first), open-source guardrail frameworks (开发者-authored), inline 安全 guards (threat-focused), AI gateways (ops-focused), and cloud-provider 护栏 (platform-native).
【机辅译·待复审】Comparison table
| Tool | Approach | 【机辅译·待复审】Indicative pricing* | Best fit | 【机辅译·待复审】Standout strength |
|---|---|---|---|---|
| AgentPolicy | 【机辅译·待复审】Policy-as-code compiler: natural-language policy → machine-checkable behavioral constraints; pre-tool/pre-output validation; versioning + audit trail; enterprise template library | 【机辅译·待复审】Free tier; Pro from ~$29/mo | 【机辅译·待复审】Teams operationalizing written policy without a platform team | 【机辅译·待复审】Starts from yourdocuments【机辅译·待复审】, not from code; audit trail built in |
| 【机辅译·待复审】NVIDIA NeMo 护栏 | 【机辅译·待复审】Open-source framework (Colang) for programmable 护栏 around LLM apps | Free (OSS) | 【机辅译·待复审】Eng teams 构建ing custom rails | 【机辅译·待复审】Deep 开发者 control; NVIDIA-backed ecosystem |
| 【机辅译·待复审】护栏 AI | 【机辅译·待复审】Open-source output validation library with reusable "guards" | 【机辅译·待复审】Free OSS; paid cloud | 【机辅译·待复审】开发者s validating structured outputs | 【机辅译·待复审】Rich guard library; CI-friendly |
| 【机辅译·待复审】Lakera Guard | 【机辅译·待复审】Inline 安全 guard: injection, jailbreak, content threats at runtime | 【机辅译·待复审】Commercial (quote) | 【机辅译·待复审】Production defense for LLM apps | 【机辅译·待复审】Strong real-time threat detection |
| Portkey | 【机辅译·待复审】AI gateway with 护栏, routing, and observability | 【机辅译·待复审】Free tier; paid gateway plans | 【机辅译·待复审】Ops teams running many LLM calls | 【机辅译·待复审】护栏 + routing + telemetry in one gateway |
| 【机辅译·待复审】WhyLabs (护栏) | 【机辅译·待复审】Data/LLM monitoring with policy checks | 【机辅译·待复审】Commercial (quote) | 【机辅译·待复审】Data-centric teams watching drift | 【机辅译·待复审】Strong observability lineage |
| 【机辅译·待复审】Amazon Bedrock 护栏 | 【机辅译·待复审】Platform-native content filters & denied topics on Bedrock | 【机辅译·待复审】Usage-based (AWS) | 【机辅译·待复审】AWS-native stacks | 【机辅译·待复审】Zero-infra integration with Bedrock apps |
【机辅译·待复审】* Indicative as of 2026; verify current pricing on vendor sites. Open-source 工具 are free to use but require engineering time.
AgentPolicy【机辅译·待复审】— deep dive
功能概览.【机辅译·待复审】AgentPolicy converts natural-language company policy — 合规 clauses, brand guidelines, privacy rules — into machine-checkable constraints that run before the agent calls 工具 or emits output. It ships multi-domain coverage (data privacy, content safety, external communications), a template library of common enterprise policies, and — critically — versioning with an audit trail: every policy change is logged and every enforcement decision is traceable to a rule version.
Pros
- 【机辅译·待复审】The input is your existing policy documents; you do not need to author a DSL first.
- 【机辅译·待复审】Behavioral layer (layer 2) is the core, not an afterthought — most alternatives start at input/output filtering.
- 【机辅译·待复审】Audit trail and versioning satisfy the "show me" moments in 合规 reviews.
- 【机辅译·待复审】Template library gives day-one coverage for the policies most companies actually have.
Cons
- 【机辅译·待复审】It enforces policy; it does not defend against adversarial 安全 threats at the network level — pair with Lakera-style runtime guards if attacker traffic is a primary concern.
- 【机辅译·待复审】Policy compilation needs a review step: automated conversion gets you 90% there, and a human signs off the rest (which is also what your auditors will want).
【机辅译·待复审】Real use case.【机辅译·待复审】A SaaS company's 合规 team had a written rule: no customer PII in any external communication 生成d by AI. AgentPolicy compiled it into an output filter plus a tool constraint (the email-sending tool strips and flags PII patterns). The first week of enforcement logs caught three near-misses that previously would have been silent — each now a documented blocked action with a rule reference.
【机辅译·待复审】Real use case (second segment).【机辅译·待复审】A consumer fintech ran support agents with brand-tone and refund-authority rules across three business lines. Central 合规 authored the rules once in AgentPolicy; each line's agent inherited them, with line-specific amounts as parameters. When the refund ceiling changed, one rule version bump updated all three agents — with the change logged for the quarterly 合规 review.
【机辅译·待复审】Choosing between the families
- 【机辅译·待复审】You have engineers who want full control:【机辅译·待复审】NeMo 护栏 or 护栏 AI — author rails in code, own everything.
- 【机辅译·待复审】Your threat model is adversarial traffic:【机辅译·待复审】Lakera Guard for inline defense, alongside policy enforcement.
- 【机辅译·待复审】You run many models and want 护栏 at the gateway:【机辅译·待复审】Portkey, or Bedrock 护栏 if you are already AWS-native.
- 【机辅译·待复审】Your starting point is written policy and 合规 evidence:【机辅译·待复审】AgentPolicy — compile what your company already believes, then enforce it.
【机辅译·待复审】Most serious stacks end up combining: a policy layer (AgentPolicy or NeMo) + a 安全 guard (Lakera or equivalent) + gateway observability (Portkey or cloud-native). The combination mirrors what app 安全 settled on: authorization + WAF + telemetry.
【机辅译·待复审】A reference policy pack (start with these five rules)
【机辅译·待复审】If you are compiling your first policy pack, these five rules cover the majority of enterprise exposure:
- PII egress:【机辅译·待复审】no personal data in external communications or third-party tool calls; pattern-detect and block.
- 【机辅译·待复审】Monetary authority:【机辅译·待复审】refund/credit/discount ceilings with human-approval gates above thresholds.
- 【机辅译·待复审】Commitments:【机辅译·待复审】the agent may not promise dates, outcomes, or 法务 positions not in the approved statements list.
- Disclosure:【机辅译·待复审】every conversational surface identifies the agent as AI (Article 50 alignment).
- Escalation:【机辅译·待复审】defined triggers (angry sentiment, 法务 keywords, competitor claims) route to a human — logged, not silent.
【机辅译·待复审】Each becomes 1–3 predicates with an owner and an audit destination. That is a complete, defensible v1 guardrail layer.
常见问题
- 【机辅译·待复审】Aren't 护栏 just better system prompts?
- 【机辅译·待复审】System prompts are instructions; 护栏 are controls. Instructions can be overridden, forgotten after context growth, and reinterpreted by model updates — and they leave no audit trail. A control evaluates deterministically, fails loudly, and logs. You need both: prompts for behavior shaping, 护栏 for boundaries you can prove.
- 【机辅译·待复审】Where should 护栏 run — in our app, at a gateway, or at the model provider?
- 【机辅译·待复审】At minimum where your agent's decisions become actions: before tool calls and before output leaves your system. Gateway-level 护栏 (Portkey, Bedrock) add uniform coverage across many calls; model-provider filters cover content classes but cannot know your refund ceiling or brand rules. Behavioral policy is yours to own.
- 【机辅译·待复审】How do we handle false positives without killing the agent's usefulness?
- 【机辅译·待复审】Tune in three moves: log-only mode first (observe what would have been blocked), narrow predicates (block the specific action, not the topic), and human-approval gates instead of hard blocks for judgment-call zones. Versioning lets you iterate rules like code — which is the entire point of policy-as-code.
- 【机辅译·待复审】Do 护栏 satisfy the 欧盟 AI 法案?
- 【机辅译·待复审】They evidence parts of it: Article 50 disclosure duties (enforceable since August 2026) map to a disclosure rule; robustness and oversight expectations for higher-risk systems map to enforcement logs and human gates. 护栏 are evidence and mechanism — they do not replace a classification and conformity process for high-risk use cases.
- 【机辅译·待复审】What's the difference between 护栏 and the 红队 testing layer?
- 【机辅译·待复审】护栏 are the fences; 红队 testing is the fence inspector. Article 2 of this series covers testing your agent for injection, escalation, and exfiltration — the mature loop is test → fix → encode the fix as policy → version it.
- 【机辅译·待复审】Can we enforce policy written by 法务/合规 people who don't code?
- 【机辅译·待复审】That is precisely the policy-compiler category. With AgentPolicy, 合规 authors bring the natural-language rules and review the compiled predicates; engineering wires the enforcement points. The review step is a feature — it keeps humans accountable for what the machine enforces.
- 【机辅译·待复审】How much latency do 护栏 add?
- 【机辅译·待复审】Predicate evaluation is typically milliseconds — negligible next to an LLM call. LLM-as-judge validators are the expensive exception; reserve them for low-volume, high-risk outputs. Design your layers so the cheap deterministic checks run on everything and the expensive ones run on the risky subset. ---
Sources
- 【机辅译·待复审】NVIDIA NeMo 护栏.【机辅译·待复审】github.com/NVIDIA/NeMo-护栏.
- 【机辅译·待复审】护栏 AI.【机辅译·待复审】github.com/护栏-ai/护栏.
- 【机辅译·待复审】EUR-Lex — Regulation (EU) 2024/1689 (欧盟 AI 法案, Art. 50).【机辅译·待复审】eur-lex.europa.eu/eli/reg/2024/1689/oj.
相关 工具
- AgentPolicy — Turn company policy into agent-enforced rules
- AgentRedTeam — Break your AI agents before attackers do
- AIActRadar — Turn EU AI Act chaos into a clear compliance roadmap