Mục lục
Nội dung này phục vụ mục đích giáo dục, không tự chứng nhận production, pháp lý, tuân thủ, đầu tư hoặc tín hiệu giao dịch.
Trả lời ngắn: Production AI System Framework là bộ tám gate giúp team quyết định một hệ thống AI có đủ bằng chứng để release và vận hành hay chưa. Framework nối product, data, architecture, evaluation, security, observability, cost, capacity và ownership thành Evidence Pack có threshold, người chịu trách nhiệm, stop condition và rollback.
Đọc xong, bạn sẽ hiểu:
- Tám gate biến 29 bài trước thành một release decision.
- Artifact, metric, fail condition và owner cần có ở từng gate.
- Cách dùng Release Readiness Card để chọn hold, canary, expand hoặc retire.
1. Production AI System Framework là gì?
Demo trả lời đúng vài câu chưa phải production system. Production phải sống cùng user thật, data thay đổi, quota, dependency, attacker, incident, cost và người trực. Cũng như máy bay không được cất cánh chỉ vì sơn đẹp, model output ấn tượng không thể thay kiểm động cơ, nhiên liệu, phi công và phương án quay đầu.
Framework này có tám gate:
- Product Contract
- Data & Trust
- Architecture & Execution
- Evaluation & Safety
- Security & Permissions
- Observability & Incident
- Cost & Capacity
- Release & Ownership
Mỗi gate phải trả lời bốn câu: artifact nào chứng minh, metric/threshold nào áp dụng, ai sở hữu, điều kiện nào buộc hold hoặc rollback. Tám gate không phải tám team ném giấy cho nhau; chúng là tám góc nhìn trên cùng một system version.
Evidence Pack là bộ link tới product contract, data card, system map, eval report, threat model, observability map, cost/load report và runbook. Artifact phải có version và ngày review. Checkbox “đã test” không cho reviewer biết test gì, trên version nào, pass theo threshold nào.
Framework không đảm bảo hệ thống an toàn tuyệt đối và không thay compliance review. Threshold tùy intended use, risk tolerance và người bị ảnh hưởng. Một trợ lý viết draft nội bộ khác hệ thống tự thực hiện giao dịch, tuyển dụng hoặc quyết định quyền lợi.
NIST AI RMF Core nhấn mạnh quy trình test, evaluation, verification và validation phải có metric, phương pháp, tài liệu; post-deployment còn cần monitoring, feedback, appeal/override, incident, recovery, change management và decommission (NIST AIRC).
Hình 1 — Production-ready là kết quả của tám gate cùng có evidence, không phải một model trả lời đẹp.
2. Gate 1 Product Contract và Gate 2 Data & Trust
Gate 1 — Product Contract
Product contract khóa problem, user, intended use, prohibited use, outcome, workflow boundary và human role. Nếu team không biết output đi vào quyết định nào, họ không thể chọn eval hay risk control đúng.
Artifact tối thiểu:
- problem statement và user;
- input/output contract;
- success và failure definition;
- side effect được phép;
- human gate, override và appeal;
- non-goal, out-of-scope, stop condition.
Metric có thể là task completion, correction rate, escalation, time saved hoặc complaint—không chỉ “accuracy”. Fail condition: không có business owner, intended use mơ hồ, hoặc output được dùng cho mục đích khác contract.
Ví dụ nhà hàng: “AI giúp draft thực đơn” khác “AI tự đặt nguyên liệu”. Cùng một model, nhưng side effect và hậu quả sai khác hẳn. Gate 1 ngăn scope creep âm thầm biến trợ lý thành người ra quyết định.
Gate 2 — Data & Trust
Gate này lập data inventory, source, owner, classification, consent/authority, retention, freshness, quality và data lineage—dấu vết dữ liệu đi từ nguồn qua biến đổi đến output. Với RAG, còn cần index version, chunking, access filter và citation path.
Trust boundary là nơi dữ liệu hoặc lệnh đi sang component có mức tin cậy/quyền khác: browser sang server, orchestrator sang model provider, model text sang tool executor. Mỗi boundary phải có validation, identity và policy phù hợp.
Artifact:
- data card và system-of-record map;
- schema/quality/freshness checks;
- access matrix và retention rule;
- third-party/provider data-flow record;
- deletion, correction và rebuild procedure.
Metric có thể là source coverage, stale-document rate, access-filter pass, citation coverage hoặc retrieval relevance. Fail condition: không rõ nguồn, dữ liệu nhạy cảm đi sai boundary, index không tái tạo được, hoặc quyền retrieval rộng hơn user.
Microsoft Azure AI architecture tách data processing, model, intelligent application và platform control; tài liệu cũng khuyên ghi nguồn retrieved trong audit trail để điều tra và giải thích quyết định sau này (Microsoft Azure).
3. Gate 3 Architecture & Execution và Gate 4 Evaluation & Safety
Gate 3 — Architecture & Execution
System map phải cho thấy entry point, orchestrator, model route, memory, retrieval, tool, queue, state store, fallback và external dependency. Vẽ cả data flow lẫn control flow; một sơ đồ chỉ có logo vendor không đủ.
Artifact:
- component/sequence diagram;
- input/output schema và version contract;
- timeout, retry budget, fallback, circuit breaker;
- state, concurrency và idempotency design;
- dependency inventory và failure-mode table.
Idempotency nghĩa là lặp cùng request không tạo side effect trùng. Nó rất quan trọng khi queue retry sau timeout: client không biết order đã tạo hay chưa, nhưng hệ thống phải biết.
Metric gồm success theo route, tool error, duplicate action, timeout, queue age và recovery. Fail condition: agent loop không giới hạn, tool action thiếu confirmation, state conflict, fallback bypass policy hoặc dependency fail làm toàn hệ thống treo.
Gate 4 — Evaluation & Safety
Eval Pack gồm dataset, scenario taxonomy, rubric, threshold, evaluator, result, failure examples và version. Test functional correctness, groundedness, format, refusal, safety, tool use, adversarial input và regression.
Average score không được che critical error—lỗi nghiêm trọng bị cấm dù điểm trung bình cao, như lộ dữ liệu, side effect sai hoặc bypass authorization. Cần tách deterministic check, model-based evaluator và human review; ghi limitation của từng loại.
Artifact:
- golden/adversarial/regression sets;
- rubric và pass threshold;
- critical-error taxonomy;
- evaluator calibration;
- error analysis và accepted residual risk;
- re-eval trigger khi model/prompt/data/tool đổi.
Metric phải theo slice: language, task type, route, user group, tool và long-tail input. Fail condition: eval không giống deployment context, test data leak vào prompt tuning, critical error >0, hoặc reviewer không tái tạo được result.
NIST AI RMF Core yêu cầu test set, metric và công cụ TEVV được tài liệu hóa; performance/assurance cần đo trong điều kiện tương tự bối cảnh triển khai và có thể cần assessor độc lập theo risk tolerance (NIST AIRC).
4. Gate 5 Security & Permissions và Gate 6 Observability & Incident
Gate 5 — Security & Permissions
Threat model liệt kê asset, actor, entry point, trust boundary, abuse case và control. Với AI, thêm prompt injection, indirect injection, data exfiltration, tool abuse, model/provider supply chain và output-to-code risk.
Artifact:
- threat model và abuse-case tests;
- identity/authorization matrix;
- secret management và egress policy;
- input/output/tool validation;
- rate limit, quota và audit logging;
- security incident contact.
Least privilege phải nằm ở tool/API, không chỉ trong prompt “đừng làm”. Model đề xuất action; deterministic code kiểm schema, permission, policy, confirmation và idempotency trước khi thực thi. Fail condition: untrusted content điều khiển privileged tool, shared secret không rotate, log chứa sensitive payload hoặc agent có quyền rộng hơn intended use.
Gate 6 — Observability & Incident
Một trace nối request ID với route, prompt/config version, retrieval IDs, model call, tool call, retry, latency, token/cost, evaluator và final disposition. Log cần đủ điều tra nhưng phải redact và retention đúng.
Observability map tách bốn lớp:
- system health: availability, error, latency, saturation;
- AI quality: groundedness, refusal, correction, drift;
- safety/security: blocked attack, policy violation, access denial;
- value/cost: adoption, successful task, cost/success.
Alert phải actionable: metric, threshold, window, severity, owner và runbook—quy trình xử lý tình huống. Runbook ghi cách xác minh, giảm blast radius, rollback, giao tiếp và thu evidence. On-call không thể xử lý alert “AI có vẻ tệ”.
Fail condition: không trace được output về version, alert không owner, incident không có kill switch, log thu quá nhiều dữ liệu hoặc rollback chưa diễn tập.
AWS khuyên production GenAI có auditability, model/prompt/application/data lineage, logging/tracing, monitoring/alerting, incident response và remediation thay vì coi governance là policy trên giấy (AWS Prescriptive Guidance).
Microsoft cũng khuyên availability và quality alert phải đủ rõ để operations hành động, đồng thời standard operating procedure phải nối operations với data/model team (Microsoft Azure Operations).
5. Gate 7 Cost & Capacity và Gate 8 Release & Ownership
Gate 7 — Cost & Capacity
Cost model gồm model, token, retrieval/data, tool, retry, eval, infra và operations. Metric chính là cost per successful task, đi cùng quality/safety/latency. Capacity plan có demand, concurrency, queue, throughput, quota và headroom—capacity còn lại trước giới hạn.
Artifact:
- cost trace theo use case/owner;
- load-test report và bottleneck map;
- budget/alert/showback rule;
- autoscaling, queue và backpressure policy;
- capacity forecast và quota dependency.
Fail condition: chưa có success denominator, load test không chạm target, critical path hết headroom, retry làm cost bùng hoặc scale bypass guardrail. “Chạy được 10 request” không chứng minh chịu 1.000 request; “autoscaling bật” không chứng minh downstream scale được.
Gate 8 — Release & Ownership
Release plan ghi environment, artifact version, migration, feature flag, canary, health gate, expansion step, halt/rollback và communication. Canary là release cho tỷ lệ nhỏ trước để giới hạn blast radius; không phải cái cớ test mù trên user.
Ownership tối thiểu:
- business/product owner;
- technical owner;
- operations/on-call;
- security/risk contact;
- approver và rollback authority;
- review/reapproval/retire date.
Safe release loop:
Baseline → Verify → Canary → Health gate → Expand
Nhánh fail:
Fail → Halt → Rollback → Review
Microsoft safe-deployment guidance nhấn mạnh progressive exposure, health check trước từng giai đoạn, phát hiện lỗi thì dừng và bắt đầu recovery (Microsoft Azure Safe Deployments).
Hình 2 — Progressive exposure chỉ mở rộng khi health gate đạt; failure phải dừng và rollback.
Sai lầm, giới hạn và release checklist
Sai lầm thứ nhất là gọi “đã có guardrail” nhưng không có test. Sai lầm thứ hai là pass average rồi bỏ critical slice. Sai lầm thứ ba là monitoring system health nhưng không đo AI quality. Sai lầm thứ tư là có rollback script chưa từng chạy. Sai lầm thứ năm là owner chỉ xuất hiện lúc ký duyệt.
Giới hạn: pre-release eval không bao phủ mọi input thật; human evaluator có bất đồng; model/provider thay đổi; attacker thích nghi; traffic và data drift. Framework giảm bất ngờ và tăng khả năng phục hồi, không xóa uncertainty.
Checklist:
- Khóa version của product, data, prompt, model, tool và policy.
- Link tám artifact vào Evidence Pack.
- Xác minh threshold/critical-error gate.
- Chạy security abuse tests và authorization denial.
- Chạy load test, cost trace và capacity headroom.
- Xem trace, alert, runbook, kill switch và incident contact.
- Diễn tập rollback ít nhất một success và một partial failure.
- Chọn canary size, health window và expansion rule.
- Ghi approver, on-call, review date và retire trigger.
Đứng ngoài nếu thiếu owner, eval version, threat model, trace hoặc rollback. Dừng khi critical error >0, authorization bypass, data boundary sai, queue mất giới hạn, health gate fail hoặc evidence không khớp artifact đang release.
6. Case PROD-30 và Release Readiness Card 15 phút
Case demo PROD-30 / REL-3001 là một assistant tìm tài liệu và draft câu trả lời; không tự gửi và không thực hiện giao dịch. Team khóa 30 scenario gồm câu thường, dữ liệu cũ, source conflict, prompt injection, tool timeout và long context.
Threshold demo:
- functional pass ≥29/30;
- grounded pass ≥28/30;
- critical error =0;
- p95 ≤5 giây;
- cost/success ≤0,15 USD;
- rollback drill 2/2;
- business owner và on-call đã gán.
Kết quả:
| Gate metric | Result | Threshold | Status |
|---|---|---|---|
| Functional | 29/30 | ≥29/30 | PASS |
| Grounded | 28/30 | ≥28/30 | PASS |
| Critical error | 0 | 0 | PASS |
| Latency p95 | 4,7s | ≤5s | PASS |
| Cost/success | $0,13 | ≤$0,15 | PASS |
| Rollback drill | 2/2 | 2/2 | PASS |
| Owner + on-call | YES | required | PASS |
Hình 3 — Mỗi gate để lại artifact có version, owner và link để reviewer kiểm được.
Hình 4 — Case qua threshold demo và rollback drill nên chỉ mở canary 10%, chưa full release.
Decision là CANARY 10%, không phải full release. Canary health gate giữ functional/grounded sample, critical error, p95, cost/success và complaint. Một critical error hoặc authorization bypass dừng ngay; p95/cost fail ba cửa sổ liên tiếp cũng halt và rollback.
Câu “7/7 metric pass nghĩa là production-safe” là đoán chắc. Câu có điều kiện: evidence hiện tại đủ mở canary giới hạn trên đúng version và intended use; team vẫn phải theo dõi traffic thật, feedback và incident.
a. Mẫu đối chiếu Release Readiness Card
Dùng Google Sheets/Docs, dữ liệu demo, không dùng secret hoặc traffic production.
| Gate | Required artifact | Metric/threshold | Current evidence | Owner | Status | Stop/rollback |
|---|---|---|---|---|---|---|
| Product | Contract v3 | intended use locked | Link P-30 | Product | PASS | scope change |
| Data | Data Card v2 | access test 100% | Link D-30 | Data | PASS | cross-tenant |
| Eval | Eval Pack v5 | critical=0 | 0/30 | QA | PASS | critical >0 |
| Security | Threat Model v2 | auth deny 12/12 | 12/12 | Security | PASS | bypass |
| Observability | Trace/Runbook | trace 30/30 | 30/30 | SRE | PASS | trace gap |
| Cost/Capacity | Load Report | p95≤5; cost≤.15 | 4,7; .13 | Platform | PASS | 3 windows |
| Release | Plan v2 | rollback 2/2 | 2/2 | Release | PASS | health fail |
| Ownership | RACI/on-call | all assigned | YES | Business | PASS | owner missing |
Trong 15 phút, điền một use case demo, dán evidence link giả và tô đỏ ô chưa có. Đừng tự “PASS” bằng cảm giác. Kết quả mong đợi là một quyết định có điều kiện: HOLD, CANARY, EXPAND hoặc RETIRE, kèm lý do và người chịu trách nhiệm.
7. Tổng kết AI MASTER SERIES
a. Năm ý chính
- Production AI là hệ thống và operating model, không chỉ model.
- Tám gate cần artifact, metric, fail condition và owner.
- Critical error không được bù bằng điểm trung bình.
- Evidence Pack phải versioned, linked và reviewable.
- Pass local/pre-release gate chỉ mở canary; production evidence tiếp tục tích lũy.
b. Câu hỏi tự kiểm tra
- Vì sao demo đẹp chưa đủ release?
- Data lineage và trace khác nhau thế nào?
- Khi nào critical error chặn release?
- Canary khác full release ở đâu?
- Evidence Pack dùng để làm gì?
c. Gợi ý đáp án
Xem gợi ý câu 1
Demo chưa chứng minh data, failure, security, operations, cost và ownership. Xem mục 1.
Xem gợi ý câu 2
Data lineage theo nguồn/biến đổi dữ liệu; trace theo một execution xuyên component. Xem mục 2 và 4.
Xem gợi ý câu 3
Khi taxonomy quy định lỗi không được phép bù bằng average; threshold thường là 0. Xem mục 3.
Xem gợi ý câu 4
Canary chỉ mở tỷ lệ nhỏ, có health gate và rollback trước khi expand. Xem mục 5.
Xem gợi ý câu 5
Nó nối evidence từng gate với version, threshold, owner và decision. Xem mục 1 và 6.
d. Thuật ngữ cần nhớ
| Thuật ngữ | Giải thích ngắn |
|---|---|
| Product contract | Problem, user, outcome và boundary đã khóa |
| Data lineage | Dấu vết nguồn và biến đổi dữ liệu |
| Trust boundary | Ranh giới đổi mức tin cậy hoặc quyền |
| Idempotency | Lặp request không tạo side effect trùng |
| Eval Pack | Dataset, rubric, threshold và result có version |
| Critical error | Lỗi không được bù bằng điểm trung bình |
| Trace | Dấu vết một execution xuyên component |
| Runbook | Quy trình xử lý tình huống vận hành |
| Headroom | Capacity còn lại trước giới hạn |
| Canary | Release cho tỷ lệ nhỏ trước |
e. Nguồn tham khảo
- NIST AIRC — AI RMF Core
- AWS — Enterprise-grade security and governance for GenAI
- Microsoft Azure — Architecture pattern for AI workloads
- Microsoft Azure — AI workload operations
- Microsoft Azure — Safe deployment practices
Điều hướng: Ôn Enterprise AI #29. Bạn đã hoàn tất chuỗi 30 bài; có thể quay về AI MASTER SERIES để review theo phase.
Nội dung này phục vụ mục đích giáo dục, không tự chứng nhận production, pháp lý, tuân thủ, đầu tư hoặc tín hiệu giao dịch. Mọi thị trường đều có rủi ro mất vốn.