technext-edge · Extractor pod

Benchmark Report — 21/09/2026

← technext-edgetechnext-edge

Ba loại bằng chứng tách biệt, không thay thế nhau: (1) unit test — code logic, dùng provider giả lập; (2) eval — model AI thật, 14 câu giả lập, có số liệu lặp lại được; (3) test tay trên WhatsApp — hội thoại thật/gần thật, không có số liệu lặp lại, nhưng là nơi tìm ra hầu hết bug thật hôm nay. Three separate kinds of evidence, none a substitute for another: (1) unit tests — code logic, mocked providers; (2) eval — real AI models, 14 synthetic messages, repeatable numbers; (3) manual WhatsApp testing — real/near-real conversations, no repeatable score, but where most of today's real bugs were actually found.

1 · Unit test (213/213 pass)1 · Unit tests (213/213 pass)

Chạy bằng npm test (Vitest), dùng provider giả lập (fakeProvider) tự trả về JSON đã ghi sẵn — kiểm tra code xử lý đúng logic không. Không gọi model AI thật, nên không kiểm tra được model có hiểu đúng câu khách không — việc đó thuộc phần 2 và 3. Run via npm test (Vitest), using a mocked provider (fakeProvider) that returns pre-written JSON — checks whether the code's own logic is correct. No real AI model is called, so this cannot verify whether a model actually understood a guest's message — that's Parts 2 and 3.

FileFile Số testTests Kiểm tra gìWhat it checks
robustness.test.ts48200 payload giả lập ngẫu nhiên (fuzz) — extract() luôn resolve/reject đúng cách, không crash200 randomly generated (fuzz) payloads — extract() always resolves/rejects correctly, never crashes
whatsapp.test.ts38Chữ ký HMAC, dedupe wamid, claim/release, xin lỗi khi lỗi, handoff, lệnh resetHMAC signature, wamid dedupe, claim/release, apology on failure, handoff, reset command
dates.test.ts31Phân giải ngày tuyệt đối/tương đối 3 ngôn ngữ, mốc TODAY cố địnhAbsolute/relative date parsing in 3 languages, fixed TODAY anchor
extract.test.ts20Retry-once, phân loại lỗi transport/validation, enforceVerbatimEvidence, postProcessRetry-once, transport-vs-validation error classification, enforceVerbatimEvidence, postProcess
questions.test.ts163 loại reply (greeting/questions/summary), nhãn (assumed), câu hỏi theo ngôn ngữ3 reply kinds (greeting/questions/summary), the (assumed) label, per-language questions
counts.test.ts14Đối chiếu số khách/đêm/phòng với văn bản gốcCorroborates guests/nights/rooms counts against the source text
evalReplay.test.ts12Bộ chấm điểm eval offline khớp đúng logic scorerOffline eval replay matches the scorer's own logic
conversationStore.test.ts9Split TTL, fencing token, pause/resume/handoffSplit TTL, fencing token, pause/resume/handoff
normalize.test.ts7Nhận diện ngôn ngữ, che PII khi logLanguage detection, PII masking for logs
converse.test.ts6Giới hạn transcript, luồng hội thoại nhiều lượtTranscript cap, multi-turn conversation flow
providers/providerFromEnv.test.ts4Chọn đúng provider theo env var, wrapper resilient fallbackCorrect provider selection by env var, resilient fallback wrapper
bff.test.ts3Route cơ bản của BFFBasic BFF routes
providers/gemini.test.ts, deepseek.test.ts4Timeout, prompt transport có đủ rule 2 ngôn ngữTimeout, transport prompt carries both languages' rules
dataset.test.ts1File dataset eval là JSON hợp lệEval dataset file is valid JSON
213/213
test passtests passing
15
file testtest files
0
lỗi typechecktypecheck errors

2 · Eval benchmark — Gemini vs DeepSeek (14 case, model AI thật)2 · Eval benchmark — Gemini vs DeepSeek (14 cases, real AI models)

Chạy qua eval/runner.mjs, gọi thật /v1/extract trên production. 14 case gồm 10 case gốc (đa ngôn ngữ, có bẫy bịa đặt) + 4 case en-06..en-09 thêm hôm 20/09 sau khi phát hiện bug guests. Run via eval/runner.mjs, calling real /v1/extract on production. 14 cases: 10 original (multilingual, with fabrication traps) + 4 cases en-06..en-09 added 20/09 after the guests bug was found.

Thời điểmWhen Provider Required fieldsRequired fields Fabrication Evidence p95 latencyp95 latency
20/09, trước khi sửa bug20/09, before the bug fixgoogle:gemini-3.1-flash-lite31/35 (89%)0100%6.5s
21/09, Gemini sau fix guests+diveTo21/09, Gemini after the guests+diveTo fixgoogle:gemini-3.1-flash-lite34-35/35 (97-100%)0100%~4s
21/09, DeepSeek ngay khi gateway sống lại (chưa qua wrapper)21/09, DeepSeek right after the gateway came back (bare, no wrapper)deepseek-gateway:deepseek-flash34/35 (97%)0100%5.2s
21/09, DeepSeek vừa lên provider chính (bug wrapper)21/09, DeepSeek just switched to primary (wrapper bug active)deepseek-gateway:deepseek-flash32/35 (91%)0100%5.0s
21/09, DeepSeek sau khi sửa hết (wrapper + checkIn cô lập)21/09, DeepSeek after every fix (wrapper + isolated checkIn)deepseek-gateway:deepseek-flash35/35 (100%)0100%5.7s
Kết luận:Conclusion: DeepSeek (provider chính hiện tại) đạt đủ cả 4 ngưỡng Playbook (fabrication=0, required≥95%, evidence=100%, p95≤8s) sau khi sửa 2 bug thật tìm được trong ngày — cả hai đều do cùng 1 cơ chế: model tự hạ độ tin cậy khi phải xử lý nhiều rule cùng lúc trong 1 lệnh gọi, sửa bằng cách tách riêng field yếu ra 1 cuộc gọi nhỏ chạy song song. DeepSeek (current primary) now clears all 4 Playbook thresholds (fabrication=0, required≥95%, evidence=100%, p95≤8s) after fixing two real bugs found today — both from the same mechanism: the model lowers its own confidence when classifying many rules in one call, fixed by isolating the weak field into its own small call run in parallel.
Lưu ý về mẫu số liệu:Note on the sample: 14 case là bộ dữ liệu giả lập (synthetic), không phải tin nhắn khách thật — dùng để rà bug và theo dõi hồi quy, không phải cơ sở quyết định model cuối cùng (quyết định thật chờ 30 tin thật từ Eloa, theo ADR-005a). These 14 cases are synthetic, not real guest messages — used to hunt bugs and track regressions, not the basis for a final model decision (that waits on Eloa's 30 real messages, per ADR-005a).

3 · Test tay trên WhatsApp — không phải benchmark, là dò tìm (exploratory)3 · Manual WhatsApp testing — not a benchmark, exploratory

Không có số liệu lặp lại được, không chạy tự động, không có tiêu chí pass/fail cố định — đây là hội thoại thật/gần thật do người gõ tay, mỗi lần đổi 1 chi tiết để dò lỗi. Giá trị: tìm ra phần lớn bug thật trong ngày, những bug mà unit test (giả lập) và eval (câu ngắn) không chạm tới. No repeatable score, not automated, no fixed pass/fail bar — these are real/near-real conversations typed by hand, changing one detail each time to probe for issues. Value: found most of today's real bugs, ones unit tests (mocked) and eval (short messages) never touched.

Kịch bảnScenario KênhChannel Kết quảResult
Hội thoại 4 lượt tiếng Việt (chào → trả lời dồn → lặn → tóm tắt)4-turn Vietnamese conversation (greeting → dense answer → dive → summary)API (/v1/converse)OK
Booking thật "Alex" — cửa sổ lặn hẹp hơn cả kỳ nghỉReal "Alex" booking — dive window narrower than the whole stayWhatsApp sandboxOK — phân biệt đúng, No-Fly advisory bắn đúngcorrectly distinguished, No-Fly advisory fired correctly
Test độ trễ 7 phút giữa 2 tin (rủi ro lớn nhất cho demo)7-minute gap test between two messages (top demo-day risk)WhatsApp sandboxOK — nhớ đúng contextcontext retained correctly
Thread "Michael" — hỏi về baby/thêm người sau khi đã cho số khách"Michael" thread — asked about a baby/extra person after guests was already givenWhatsApp sandboxTìm ra 2 bugFound 2 bugs — nhãn (assumed) mất, guests bị xoá về missing(assumed) tag lost, guests wiped to missing
Email dài kiểu gia đình có con (Jennifer) — số tổng + chi tiết lặn của 1 ngườiLong family-style email (Jennifer) — a stated total plus one person's diving detailAPI + WhatsApp sandboxTìm ra 2 bugFound 2 bugs — guests, diveTo (cả hai đã sửaboth since fixed)
OK = chạy đúng như kỳ vọngbehaved as expected Tìm ra bugFound bug(s) = test tay bắt được lỗi mà unit test/eval không bắt đượcmanual testing caught something unit tests/eval missed

4 · Bug tìm được & đã sửa trong ngày 21/094 · Bugs found & fixed on 21/09

Bug Nguyên nhânRoot cause Cách sửaFix Trạng tháiStatus
guests bị bỏ vềwiped to missingModel quá tải khi phân loại nhiều field cùng lúc trong 1 promptModel overloaded classifying many fields at once in one promptCô lập guests ra 1 cuộc gọi riêng, chạy song songIsolated guests into its own call, run in parallelĐã sửaFixed
diveTo bỏ sót ngày thứ 2 ("16th and 17th")missed the second date ("16th and 17th")Thiếu rule cho cú pháp "ngày X and ngày Y"Missing rule for the "date X and date Y" patternThêm rule + ví dụ vào prompt chínhAdded rule + example to the main promptĐã sửaFixed
extractGuests không chạy trên productiondidn't run in productionWrapper resilient (Gemini↔DeepSeek fallback) không forward optional methodResilient (Gemini↔DeepSeek fallback) wrapper didn't forward the optional methodSửa wrapper forward đúngFixed the wrapper to forward extractGuests/extractCheckInĐã sửaFixed
checkIn ngày tương đối tiếng Anh không ổn định trên DeepSeekEnglish relative-date unstable on DeepSeekCùng cơ chế quá tải như bug guestsSame overload mechanism as the guests bugCô lập checkIn ra 1 cuộc gọi riêngIsolated checkIn into its own callĐã sửaFixed
Fabrication ngày ("cuối tháng này" → bịa ngày cụ thể)Date fabrication ("end of this month" → invented a specific day)Phát sinh khi viết prompt cô lập lần đầu, chưa dạy tôn trọng "chưa chốt ngày"Introduced while first writing the isolated prompt, which hadn't yet been taught to respect "not decided yet"Bắt được trước khi deploy, sửa prompt trước khi lên productionCaught before deploying, fixed the prompt before it reached productionĐã sửa, chưa từng lên productionFixed, never reached production