← Return to HomeAI Engineering Lab
HIGH-CONCURRENCY UNIFIED CORE · SMALL LAB ON-PREMISE BENCHMARK · ENTERPRISE AI LAB

Deep-Dive: Core Engineering & Multi-Level Verified Architecture

How we achieved sub-second latency (297ms P50), 52 req/s throughput, and zero-hallucination statutory accuracy on a single commodity lab node through Heterogeneous Pipelined Parallelism and Multi-Tier Data Verification.

FAST-PATH LATENCY
297 ms
P50 Sub-second Fast-Path
CONCURRENT THROUGHPUT
52.0 req/s
Single commodity node
BOUNDS SAFETY
100% Clean
4.5M+ byte offsets
LEGAL GRAPH
480K+ Edges
Active validity relations
Root Problems & System Philosophy

3 Fatal Bottlenecks in Legal Retrieval & Core System Philosophy

Why this system was engineered: Overcoming the failures of both legacy keyword search and mainstream AI models via deterministic, verifiable legal infrastructure.

COLLOQUIAL GAPPOINT 01

Colloquial vs Statutory Mismatch

Citizens search using everyday natural phrasing ('unjustly fired', 'stolen insurance'). Legacy keyword engines match raw strings and return zero results because they demand exact formal statutory terms.

Result: Zero results or irrelevant noise
HIERARCHY & VALIDITYPOINT 02

The Expired Instrument Trap

Repealed decrees still match search keywords. Plain keyword lookups and static AI scans continue to cite obsolete laws because they lack temporal validity hierarchy trees.

Result: Citing invalid, repealed laws
AI HALLUCINATIONPOINT 03

Mainstream AI Hallucinations

Mainstream AI models fabricate plausible-sounding Article and Clause numbers that do not exist in statutory reality, lacking byte-level character offsets for official Gazette verification.

Result: Fabricated legal citations
UNIFIED ARCHITECTURAL SOLUTION
CORE SYSTEM PHILOSOPHY

Deterministic Empirical Legal Computation Infrastructure

ZERO HALLUCINATION

Rather than relying on primitive keyword lookups or loose, hallucination-prone AI wrappers, our system is engineered as Deterministic Legal Infrastructure: Resolving all 3 core bottlenecks — accurately understanding citizen natural language, enforcing real-time legal validity, and eradicating hallucinations via byte-level Gazette verification.

01. 3NF SEMANTIC BRIDGE

Relational synonym normalization allowing citizens to search with everyday phrasing and match exact statutory articles without mastering legalese.

02. REAL-TIME VALIDITY GRAPH

Parent-child legal relationship modeling across timelines to automatically exclude repealed or superseded clauses in real-time.

03. BYTE-LEVEL PROVENANCE

Eliminates AI hallucination entirely by anchoring every citation strictly to character coordinates in original Gazette texts for 1-click verification.

EMPIRICAL ENGINEERING · KEY CAPABILITY LEAPS

Architecture Breakthroughs: Continuous Capability Leaps

System growth is not measured by calendar days, but by concrete technical breakthroughs that systematically resolved fundamental bottlenecks in Vietnamese legal AI.

PHASE01/07
Milestone NavigationPhase #1
Semantic Retrieval & Ground-Truth Baseline across 70,000+ Core Statutory Instruments
#1 · CORE STATUTORY CORPUS & SEMANTIC GROUNDINGPhase 01 / 07
Stack: Python ML Prototyping + 1,875 Ground-Truth Legal Dispute Benchmark

Semantic Retrieval & Ground-Truth Baseline across 70,000+ Core Statutory Instruments

Core Legal Problem:The core central statutory corpus (Labor & Social Insurance, Corporate & Tax, Land & Construction, Civil & Criminal Codes) governs the vast majority of daily citizen and business legal disputes. Naive keyword search failed on informal citizen phrasing ('getting fired', 'unpaid insurance'), while commercial LLMs fabricated non-existent article numbers without evidentiary backing.
Engineering Advances:
  • ✓Curated and standardized 70,000+ highest-impact national statutory instruments (Codes, Laws, Decrees, Circulars) into structured Article/Clause AST hierarchies
  • ✓Curated 1,875 real-world legal dispute questions verified by legal experts to establish the system's ground-truth baseline
  • ✓Applied dense vector embeddings to bridge informal citizen queries directly into statutory provisions, eliminating ungrounded LLM hallucinations (Top-1: ~40%, Top-10: ~80%)
Schematic#01
Semantic Vector Space & AST70,000+ Docs
Vector Cosine Projection (2D PCA)Similarity: 0.89
Query Probe:
"Người lao động bị đuổi việc vô cớ"
→ BLLĐ 2019 (Điều 36: Chấm dứt HĐLĐ)cos: 0.89
→ Luật BHXH (Điều 216: Trợ cấp thôi việc)cos: 0.74
→ Luật Đất đai (Bồi thường giải tỏa)cos: 0.12
Hierarchical AST Parser:
Văn bản→Chương→Mục→Điều→Khoản
Accuracy
Top-1: ~40% · Top-10: ~80%
Throughput
1.54 req/s (12.5s avg)
Scale
1,875 Ground-Truth Disputes (70K+ Docs)
Outcome: Verified semantic search viability on core statutory law; exposed the critical risk of citing expired/repealed decrees.
EMPIRICAL BENCHMARK COCKPIT · DETERMINISTIC LEGAL CORE

Đo Đạc Hiệu Năng & Độ Chính Xác Thực Chứng

Kết quả đo đạc thực nghiệm từ các đợt kiểm chuẩn chất lượng và tải đa luồng trên 01 node máy chủ lab on-premise (1x CPU + GPU).

Bước Xử Lý Trong PipelineTrung Bình (Mean)Trung Vị (P50)P90P95Min / Max
1A1. Strict FTS Search (PostgreSQL GIN)
45.04 ms39.00 ms62.00 ms68.55 ms15.8 / 78.0 ms
1A2. Postgres FTS Tuple Hydration
1.20 ms1.00 ms2.10 ms2.80 ms0.0 / 3.5 ms
1B. Qdrant Dense Vector Search (HNSW Cosine)
90.64 ms88.50 ms110.00 ms122.48 ms14.1 / 135.0 ms
2. Postgres Unit Metadata Hydration
3.50 ms3.00 ms6.00 ms8.50 ms1.0 / 9.5 ms
4. Graph Legal Relation Expansion & Assembly
156.88 ms165.76 ms310.00 ms333.84 ms110.0 / 360.0 ms
TỔNG ĐỘ TRỄ FAST-PATH (SUB-300MS MEDIAN)SUB-SECOND352.90 ms297.26 ms495.00 ms536.17 ms240.0 / 718.0 ms
Tùy chọn: Deep AI Cross-Encoder Reranker (BGE-Reranker-v2-m3)
4,848.40 ms4,859.00 ms4,939.40 ms4,955.00 ms4,713.0 / 4,970.0 ms
Độ Trễ Trung Vị (P50)
297.26 ms (~0.3 giây)

Phản hồi chớp mắt trên 100% các câu hỏi thường gặp.

Giới Hạn Đuôi Latency
0.00% Câu Quá 3 Giây

100% câu hỏi đều phản hồi dưới 0.79s (Max: 718ms).

Thông Lượng Batch Write
52.0 req/s (Tăng 22.6x)

Xử lý 100 requests batch update chỉ trong 1.92 giây.

MODULE 01 · SYSTEM DESIGN

Heterogeneous Pipelined Parallelism

Traditional RAG pipelines execute sequentially, wasting 90% of hardware capacity during I/O waits. Our Unified High-Concurrency Core uses virtual threads and reactive pipelines to interleave workloads across CPU, PostgreSQL FTS, Qdrant Vector DB, and GPU Rerankers with zero idle cycles.

Workload Latency Overlapping Schedule (No Hardware Idle)Virtual Threads + Reactive Pipeline
Host CPU (Virtual)
Requests R1..R3 (Parser)
Requests R4..R6 (Normalize)
Requests R7..R10 (AST Map)
PostgreSQL FTS
FTS Batch R4..R5 (20.5ms)
FTS Batch R7..R8
FTS Batch R1..R3
Qdrant Vector
Vector Cosine R7..R8
Vector Cosine R1..R2
Vector Cosine R4..R6
GPU Reranker
BGE-Reranker R9..R10
BGE-Reranker R3..R4
BGE-Reranker R5..R7
0%
Hardware Idle Time
100% capacity interleaved
52.0 req/s
Concurrent Throughput
Single commodity node
131.2x
Batch Update Speed
From 5,090ms to 38.8ms/batch
MODULE 02 · DATA INTEGRITY

Multi-Level Data Verification & Provenance Pipeline

Every legal entity and relation edge is governed by a strict hierarchical confirmation level (Levels 1 to 4) where higher-order ground truth automatically overrides lower-level extractions with 100% bounds safety.

L1

Level 1: Automated High-Speed Ingestion

High-speed syntax boundary and entity extraction across 70,000+ core statutory instruments.

70,000+ Core Instruments
L2

Level 2: Autonomous Multi-Agent Cross-Audit

Autonomous AI agent network continuously reading, cross-auditing, and reconciling semantic coherence against authoritative government legal portals.

Autonomous Multi-Agent Audit
L3

Level 3: Expert Human-in-the-Loop Precedence

Judicial paralegals and domain experts reviewing and certifying complex regulatory disputes with non-destructive override authority.

Highest Priority Precedence
L4

Level 4: Golden Benchmark Quality Gate

Expert-certified golden ground-truth benchmark suite serving as the immutable quality gate for the entire system.

Golden Benchmark Gate
100% Bounds Clean Guarantee: 4,517,833 byte-level character offsets verified with 100% Bounds Clean fidelity (0 negative offsets, 0 inverted boundaries, 0 empty snippets).
MODULE 03 · STRUCTURED EXTRACTION

Dedicated Tabular Structuring & Multi-Page OCR Smart Stitching

Vietnamese statutory appendices contain complex multi-page tables (customs tariffs, land price brackets across 63 provinces, engineering quotas, and civil service salary scales). Our engine reconstructs them into dedicated structured query stores.

01. SMART STITCHING

Multi-Page OCR Stitching

Multi-Page OCR Stitching: Automatically resolves broken table headers and column alignment across scanned PDF page boundaries.

02. ZERO-TABLE MECHANISM

Zero Missing Table Cells

Zero-Table Elimination: Reclaims inline narrative tables that fail standard markdown table converters.

03. DEDICATED SCHEMAS

Tariffs & Land Price Brackets

Dedicated Tabular Schemas: High-speed structured queries on land valuation and import tariff schedules.

MODULE 04 · SEMANTIC GROUNDING

3NF Legal Thesaurus & Colloquial-to-Statutory Mapping

Citizens and business executives often speak in colloquial or informal terms. Our 3NF Thesaurus Engine maps colloquial slang into exact legal provisions without hallucination.

"trốn đóng bảo hiểm" (Everyday colloquial)
Article 216 (Penal Code & Social Insurance Law) — "Evasion of social or health insurance contributions"
"bị đuổi việc" (Informal complaint)
Article 36 (Labor Code 2019) — "Unilateral termination of employment contract"
"làm thêm giờ không lương" (Workplace dispute)
Article 98 & 107 (Labor Code 2019) — "Overtime wage calculation & maximum hours"