Skip to main content
reopt Handbook
reopt Handbook
LLMOps and AgentOps in Production

Architecture and Release

Ch1. System ArchitectureCh2. Versioning and ReleaseCh3. Evaluation Framework

Operational Reliability

Ch4. Online GuardrailsCh5. Observability and SLOsCh6. Cost and Latency Optimization

Growth and Response

Ch7. Experiment OperationsCh8. Incident Management Runbook

Verification

Verification ReportVerification ArchiveUpdates
Handbook›LLMOps and AgentOps›Verification Report
한국어English

Verification Report

Structure, link, metric, and logic verification for LLMOps and AgentOps in Production

Verification Baseline

2026-05-18 (English edition translated from the Korean 4th verification baseline dated 2026-05-17)

Verification Scope

  • Page structure: meta.json declarations match actual files.
  • Formula and metric consistency: Unit cost, Error budget, Burn rate, SLI/SLO.
  • Cross-chapter flow: evaluation to release to observability to incident response.
  • External reference link reachability.

History Archive

2026-03-13 and 2026-03-26 verification records are separated into Verification Archive. This report focuses on the latest operating baseline and 4th verification.

Structure Verification

ItemResult
meta.json pages12
MDX file count12
Internal link errors0
Missing/duplicate chapters0

Logic Consistency

CheckStandardResult
Cost formula consistencyUnit cost definition is consistent between Index and Ch6Pass
Gate linkageCh3 evaluation criteria feed into Ch2 release gatesPass
SLO to incident linkageCh5 Error budget links to Ch8 incident classificationPass
Experiment safetyCh7 decision formula does not conflict with Ch4 guardrailsPass

Counterexample Scenarios

ScenarioExpected BehaviorResult
Model upgrade improves quality by +2% but raises cost by +20%Cost gate holds releasePass
Latency is normal but policy violation rate risesSafety gate blocks firstPass
SLO passes overall but a specific tenant fails more oftenTenant-segment metrics detect anomalyPass

External Link Check

CategoryLinkStatus
Google SRE Workbookhttps://sre.google/workbook/table-of-contents/Checked
OpenAI Agents SDKhttps://developers.openai.com/api/docs/guides/agentsChecked
OpenAI Pricinghttps://openai.com/api/pricing/Checked
MCP Specificationhttps://modelcontextprotocol.io/specification/2025-11-25Checked
A2A Specificationhttps://a2a-protocol.org/latest/specification/Checked
OpenTelemetry GenAIhttps://opentelemetry.io/docs/specs/semconv/gen-ai/Checked
OWASP AOShttps://aos.owasp.org/aos/Checked
OWASP MCP Top 10https://owasp.org/www-project-mcp-top-10/Checked
OWASP Agentic Skills Top 10https://owasp.org/www-project-agentic-skills-top-10/Checked
Claude Guardrailshttps://docs.claude.com/en/docs/test-and-evaluate/strengthen-guardrails/handle-streaming-refusalsChecked
Anthropic Pricinghttps://platform.claude.com/docs/en/about-claude/pricingChecked
DeepSeek Pricinghttps://api-docs.deepseek.com/quick_start/pricing/Checked
LangSmith Fleethttps://www.langchain.com/blog/introducing-langsmith-fleetChecked
Braintrust Loophttps://www.braintrust.dev/docs/loopChecked
PagerDuty AI Ecosystemhttps://www.pagerduty.com/newsroom/pagerduty-expands-ai-ecosystem-to-supercharge-ai-agents/Checked

4th Verification Details

Freshness Corrections

ItemVerificationResult
A2A latestOfficial specification identifies latest released version as v1.0.0Corrected in text
MCP 2025-11-25OAuth 2.1, Protected Resource Metadata, Client ID Metadata Documents, token audience binding, token passthrough prohibition verifiedCh1 expanded
OpenAI pricingGPT-5.5, GPT-5.4, GPT-5.4 mini pricing verified. Removed older GPT-5.4 nano framingCh6 corrected
Anthropic pricingOpus 4.7/4.6/4.5, Sonnet 4.6/4.5, Haiku 4.5 pricing and prompt caching multiplier verifiedCh6 corrected
DeepSeek pricingCurrent official pricing centers on DeepSeek V4 Flash/Pro. Removed V3.2-centered pricing tableCh6 corrected
OpenTelemetry GenAIGenAI semantic conventions status verified as DevelopmentCh5 corrected
OWASP AOSAOS verified as work-in-progress public projectCh5 corrected
OpenAI Agents SDKGuardrails, human review, resumable state, MCP, tracing, agent evals, and voice-agent operating surfaces verifiedCh3-Ch5 expanded
OWASP MCP/SkillsMCP Top 10 and Agentic Skills Top 10 controls for supply chain, permissions, and telemetry verifiedCh4/Ch8 expanded
Link statusLangSmith Fleet, Braintrust Loop, and PagerDuty source links replaced with current official URLsUpdates corrected

4th Verification Sources

SourceChecked Area
OpenAI API PricingGPT-5.5/GPT-5.4/GPT-5.4 mini, Batch, tool/container pricing
OpenAI Agents SDK docsguardrails, human review, MCP, tracing, agent evals, voice agents
Model Context Protocolcurrent specification 2025-11-25, authorization security
A2A Protocollatest v1.0.0, task/streaming/push notification/security considerations
OpenTelemetryGenAI semantic conventions Development status
OWASPAOS, MCP Top 10, Agentic Skills Top 10
Anthropic Claude docsmodel pricing, prompt caching, long context pricing
Claude guardrails docsstreaming refusal handling
DeepSeek API Docscurrent models/pricing and V3.2 release context
LangChain/Braintrust/PagerDutyFleet, Loop, AI operations ecosystem

Verification Limits

Scope

This verification focuses on document structure and operating-framework consistency. Vendor features and API signatures can change. Recheck model prices, discounts, deprecations, and benchmark names against official pages before a release.

Related docs

Updates

Changelog for LLMOps and AgentOps in Production

Verification Archive

Historical verification records and correction history for LLMOps and AgentOps in Production

Updates

Harness Engineering · Change log for the Harness Engineering handbook

Verification Report

Claude Code Complete Guide · Verification checklist for the English Claude Code handbook locale.

Earlier verification records

Claude Code Command Master · Claude Code: Earlier review dates, version baselines, and findings, with a link to the current verification scope.

Ch8. Incident Management Runbook

Operate a unified standard for quality regressions, cost spikes, and policy bypass incidents

Verification Archive

Historical verification records and correction history for LLMOps and AgentOps in Production

On this page

Verification ScopeStructure VerificationLogic ConsistencyCounterexample ScenariosExternal Link Check4th Verification DetailsFreshness Corrections4th Verification SourcesVerification Limits