Đi đến nội dung chính
D2 Group
← Automation Work

Automation case study · Document Operations

AI có thể đọc document. Workflow vẫn phải quyết định output đó có đủ an toàn để dùng hay chưa.

Case này là Architecture prototype tách AI interpretation khỏi business acceptance. Extraction, validation, classification, anomaly handling, PII protection, durable state và review/recovery paths được thiết kế thành các control layers riêng trước khi structured data trở thành accepted record.

Proof summary

Đọc evidence trước khi đọc outcome.

Evidence type

Workflow architecture + explicit evidence boundary

Evidence status

Architecture prototype

Measurement / operating scope

Structured extraction, deterministic validation, anomaly handling, PII redaction and observable exception paths

Observed state

The architecture demonstrates how AI extraction is bounded by deterministic validation and exception handling.

Claim boundary

No extraction-accuracy, error-reduction, production-volume, processing-time or cost-savings metric is claimed without a defined evaluation set and telemetry.

Proof reviewed

2026-09-25

Direct answer

Architecture này giải quyết vấn đề gì?

Nó ngăn một model response bị nhầm thành authoritative business record. Model tạo candidate interpretation; deterministic controls kiểm identity, completeness, business rules và anomalies; sensitive data có boundary riêng; cuối cùng document phải đi vào một durable state có thể audit, review hoặc recover.

Document control pipeline

Intake → Extract → Validate → Classify → Detect anomaly → Protect → Persist & decide → Observe & recover.

AI là interpretation layer. Acceptance, side effects và recovery vẫn thuộc workflow controls có thể kiểm tra.

01

Intake

Nhận document tại boundary rõ, tạo document identity, source context và processing version trước khi extraction bắt đầu để mọi output phía sau còn truy ngược được về file/source ban đầu.

02

Extract

Dùng AI-assisted extraction — trích xuất có AI hỗ trợ — để chuyển nội dung không cấu trúc thành một output shape đã định nghĩa; model output vẫn chỉ là candidate structured data.

03

Validate

Kiểm schema, required fields, data types và deterministic business rules. Một JSON đúng cấu trúc vẫn có thể sai nghĩa business và phải bị chặn trước downstream side effect.

04

Classify

Gán document type hoặc operational category bằng bounded logic; classification hỗ trợ routing nhưng không mặc định cấp quyền approve, post, pay hoặc ghi đè source-of-truth record.

05

Detect anomaly

Đưa missing fields, impossible values, cross-field conflicts, low-evidence outputs hoặc unsupported document patterns vào explicit exception state thay vì ép thành success.

06

Protect

Redact, minimize hoặc restrict PII — dữ liệu định danh cá nhân — trước persistence hoặc downstream exposure theo scope; sensitive fields không nên được copy rộng chỉ vì model đã đọc được.

07

Persist & decide

Lưu accepted structured record, processing metadata và status hoặc chuyển sang Review / Reject / Retry; raw/model output và authoritative business record phải là hai lớp khác nhau.

08

Observe & recover

Giữ processing state, exception reason, retries và recovery ownership đủ rõ để failed documents không biến mất sau một workflow run màu đỏ hoặc xanh.

Control model

Model đọc nội dung. Workflow sở hữu acceptance boundary.

01

AI interprets; deterministic controls accept

Model giúp đọc và diễn giải nội dung. Required fields, business rules, allowed states, routing và side effects vẫn phải do inspectable workflow logic quyết định.

02

Extraction confidence không phải calibrated correctness

Model confidence hoặc heuristic score có thể dùng như một signal để route review, nhưng không nên tự được diễn giải thành xác suất đã hiệu chuẩn rằng field hay record chắc chắn đúng.

03

Schema-valid không bằng business-valid

Một record có đủ field và đúng type vẫn có thể sai currency, date relation, ownership, duplicate identity hoặc cross-field rule; business validation phải là control riêng.

04

Classification không phải authorization

Gắn nhãn Invoice, Contract hay Application không đồng nghĩa workflow được phép tự động thanh toán, ký, phê duyệt hoặc ghi đè system-of-record mà không có authority rule riêng.

05

PII protection là visible system layer

Sensitive-data handling phải được biểu diễn thành redaction, minimization, access/persistence decision rõ; không nên nằm ẩn trong prompt hoặc assumption của model.

06

Every document has durable state

Accepted, Review, Reject, Retry và Failed cần có durable status, reason và owner để operator biết document đang ở đâu sau retries, restarts hoặc manual review.

Published evidence

Public case thực sự chứng minh điều gì?

Evidence chứng minh architecture và control/failure paths. Nó không support việc tự thêm accuracy %, throughput, ROI hoặc compliance certification.

01

Workflow architecture

Public case chứng minh staged document-processing architecture thay vì một LLM call duy nhất, với control/failure boundaries visible trong flow.

02

Structured extraction

Document content được chuyển sang explicit structured shape để downstream validation có contract cụ thể thay vì nhận free-form model text.

03

Deterministic classification and routing

Operational routing được biểu diễn thành inspectable workflow logic; AI interpretation không sở hữu toàn bộ end-to-end decision.

04

Anomaly and exception handling

Unexpected values, missing evidence và invalid combinations có explicit exception path thay vì bị che bởi successful model response.

05

PII protection layer

Sensitive-data handling được đặt thành dedicated architecture layer trước persistence/downstream distribution trong scope công bố.

06

Operational logging and state

Processing status và exceptions là một phần của architecture để green execution không phải evidence duy nhất cho business success.

Document outcome states

“Processed” chưa đủ. Document cần một business state rõ.

01

Accept

Record đã pass structural, business và required evidence checks trong scope; accepted record có thể tiếp tục sang approved downstream step.

02

Review

Document đủ để hiểu sơ bộ nhưng có uncertainty, anomaly hoặc missing evidence; item phải giữ reason, source context và owner cho human/specialist review.

03

Reject

Document không đạt minimum acceptance rules hoặc không thuộc supported scope; workflow không được biến nó thành authoritative business record.

04

Retry / Recover

Failure nằm ở extraction service, downstream API, persistence hoặc transient processing; retry/recovery diễn ra từ known state thay vì silently dropping document.

Identity & state

Raw document, candidate extraction và accepted record không phải một object.

Source document identity

ID/file hash/source key dùng để biết document nào thực sự đang được xử lý và tránh coi cùng source upload lại như một business record mới khi scope không cho phép.

Processing version

Version của extraction schema, prompt/model configuration hoặc deterministic rules cần đủ để giải thích vì sao cùng document ở hai thời điểm có thể tạo output khác nhau.

Candidate extraction

Model-generated structured data trước khi deterministic validation và acceptance; candidate không phải system-of-record value mặc định.

Accepted record

Structured data đã pass controls được định nghĩa cho scope và được persist với source/processing context; acceptance vẫn không vượt quá authority của workflow.

Exception identity

Stable reference cho một failed/review item để retry, human correction và audit không tạo thêm một document state rời rạc khó reconcile.

Human review model

Human-in-the-loop cần reason, context, owner và resolution state.

Reason

Review phải nói rõ vì sao: low evidence, missing field, anomaly, unsupported type, identity conflict hoặc downstream failure; không chỉ gắn nhãn Manual Review chung chung.

Evidence context

Reviewer cần thấy source document, candidate fields, validation failures và relevant processing metadata để quyết định mà không phải dựng lại flow từ đầu.

Owner

Exception phải thuộc một role/queue/team rõ; human-in-the-loop chỉ có ý nghĩa khi có ownership và trạng thái pending có thể quan sát.

Resolution

Approve, correct, reject hoặc request-more-evidence phải tạo một deterministic state transition có audit context, không phải chỉnh dữ liệu bên ngoài workflow rồi bỏ mất lịch sử.

Feedback boundary

Human correction có thể trở thành evaluation/training signal nếu governance cho phép; correction không nên tự động thay model behavior hoặc source truth mà không có process riêng.

Claim boundary

Architecture evidence không phải accuracy, compliance hay production-performance evidence.

Architecture prototype ≠ production deployment

Status public là Architecture prototype với workflow architecture + explicit evidence boundary. Case không claim production volume, latency, uptime, ROI hoặc straight-through-processing rate.

AI extraction ≠ source truth

Model output là candidate interpretation. Authoritative value cần validation và business/source authority phù hợp trước khi downstream system được phép coi nó là truth.

Confidence score ≠ accuracy percentage

Model/heuristic confidence không tự động là calibrated accuracy metric; case không công bố measured field-level hoặc document-level accuracy.

Schema-valid ≠ business-valid

Đúng JSON/type không chứng minh record hợp lệ về domain; business constraints, identity, relationships và authority vẫn phải kiểm riêng.

Classification ≠ authorization

Document type/category không cấp quyền cho consequential side effects. Payment, approval, contract state hoặc system-of-record writes cần authority rule riêng.

PII redaction ≠ compliance certification

Redaction/minimization là technical controls. Legal basis, retention, consent, access, residency và regulatory obligations vẫn cần environment-specific governance.

Green workflow ≠ accepted document

Workflow có thể chạy xong nhưng document vẫn ở Review, Reject hoặc downstream-pending state; accepted business outcome cần durable status/evidence riêng.

FAQ

AI Document Intelligence — câu hỏi thường gặp.

AI Document Intelligence Pipeline trong case này làm gì?

Đây là Architecture prototype cho document operations: nhận document, AI-assisted extraction, deterministic validation, classification, anomaly handling, PII protection, durable state và explicit review/recovery paths trước khi structured data được coi là usable.

Case này đã là production deployment chưa?

Chưa. Evidence public là workflow architecture + explicit evidence boundary; status là Architecture prototype. Page không claim production throughput, latency, uptime, ROI, accuracy hoặc straight-through-processing rate.

AI được dùng ở đâu trong workflow?

AI được dùng ở các bước cần interpretation như structured extraction hoặc semantic classification. Required fields, business rules, routing, persistence, exception handling và consequential side effects vẫn cần explicit controls.

Output từ AI có được ghi thẳng vào system of record không?

Không nên mặc định như vậy. Model output là candidate extraction; workflow cần validation, authority và accepted-state rules phù hợp trước khi một value được ghi vào business system như authoritative record.

Confidence score có phải xác suất field đúng không?

Không mặc định. Score phụ thuộc model/heuristic và chỉ có ý nghĩa calibrated accuracy nếu có evaluation/calibration riêng. Case này không công bố field-level hay document-level accuracy percentage.

Schema validation có đủ để accept document không?

Không. Schema validation chỉ kiểm structure/type. Business validation còn phải kiểm identity, date/currency relationships, allowed states, ownership, duplicate conditions và domain rules tùy use case.

Classification có thể tự động approve hoặc thanh toán không?

Classification chỉ gán document category. Approval, payment, contract state hoặc các consequential side effects cần authority rules và validation riêng; category không phải authorization.

Document nào nên đi human review?

Các document có low evidence, missing business-critical fields, anomaly, identity conflict, unsupported type hoặc consequential decision vượt automation boundary nên đi Review với reason, owner và source context rõ.

PII redaction có nghĩa hệ thống đã compliant chưa?

Không. Redaction/minimization là technical control. Compliance còn phụ thuộc legal basis, consent, retention, access, data residency, processor/subprocessor rules và regulation theo environment thực tế.

Tại sao cần document identity và processing version?

Document identity giúp chống duplicate/mất lineage; processing version giúp giải thích output được tạo bằng extraction schema, rules hoặc model configuration nào, đặc biệt khi logic thay đổi theo thời gian.

n8n đóng vai trò gì trong architecture này?

n8n có thể làm orchestration layer cho intake, extraction calls, deterministic validation, routing, persistence, review và recovery. Model service, document store và authoritative business systems vẫn là các components riêng.

Muốn đưa architecture này lên production cần thêm gì?

Cần environment-specific security/access controls, document/source contracts, evaluation set, acceptance thresholds, idempotency, durable state, exception ownership, observability, downstream verification, recovery tests và privacy/compliance governance phù hợp.

Nếu document workflow hiện tại dừng ở “AI đã extract xong”, production gap vẫn còn rất lớn.

Hãy mang document types, source systems, required fields, acceptance rules, privacy constraints và downstream actions. D2 có thể scope extraction, validation, review và recovery boundaries trước khi automation được mở rộng.

Trao đổi scope