Model summary
- Base model:
Qwen/Qwen3-0.6B
- Adapter type: PEFT LoRA
- Task: Vietnamese abstractive legal summarization
- Language: Vietnamese
- LoRA rank: 16
- LoRA alpha: 32
- LoRA dropout: 0.05
- Target modules:
q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
- Training examples in the final CEFC-RFT update: 2,394
- Reward-selection sources: 5,000
- Fine-tuning: 1 epoch
- Learning rate:
2e-5
- Gradient accumulation: 16
- Reported random seed for reward selection: 42
This repository contains an adapter, not the full Qwen3-0.6B weights. The base model is downloaded separately at inference time.
Method
For each legal source document, CEFC-RFT uses an exploit-first conditional candidate policy:
- Stage 1 — deterministic exploitation: standard legal summarization prompt + greedy decoding.
- Stage 2 — directed deterministic exploration: a legal-action-focused prompt is invoked only when the current candidate fails the required gates.
- Stage 3 — bounded stochastic exploration: a sampled candidate is generated only for still-unresolved documents.
Candidates are evaluated using source-grounded faithfulness, coherence, relevance, and deterministic legal-metadata checks. Feasible candidates are prioritized before relevance-coherence utility is optimized.
RL-lite-500 is an internal developmental pilot baseline, not a previously published method and not a pure ablation. It uses the same SFT starting point but always constructs a fixed three-candidate set (standard greedy, legal-action greedy, and one sampled candidate) and ranks it with a scalar quality-minus-penalty score before pseudo-label fine-tuning. It is retained because it exposed the motivating trade-off: high faithfulness but conservative relevance/coverage and three generations per source. CEFC-RFT changes both candidate allocation and candidate selection, so the pilot-to-final comparison should not be interpreted as a scale-only ablation.
In the reported 5,000-source run:
- 12,936 candidates were generated instead of 15,000 under fixed
K=3.
- This corresponds to a 13.76% reduction in candidate generation.
- 2,395 selected candidates were feasible.
- 2,394 were retained for weighted pseudo-label fine-tuning.
- 2,176 / 2,394 (90.89%) retained pseudo-labels had zero metadata cost.
The exported run configuration is included in cefc_run_config.json.
Reward evaluator used by CEFC-RFT
Candidate scoring in the reported CEFC-RFT run uses a frozen EViSE-mmBERT evaluator:
- Evaluator checkpoint:
quancute/mmevalsumviet2-mmbert-2048
- Backbone:
jhu-clsp/mmBERT-base
- Input: joint source-document / candidate-summary pair
- Joint input budget: 2,048 tokens
- Outputs: separate bounded scores for Faithfulness (F), Coherence (C), and Relevance (R)
- Representation: CLS + masked-mean pooling
- Training objective: criterion regression + within-document pairwise ranking
- Reference-free at inference: the evaluator needs the source and candidate summary, not a gold summary.
EViSE-mmBERT is used only as a proxy reward evaluator during CEFC-RFT candidate gating and selection; its parameters are not updated by CEFC-RFT. Its published model-card evaluation is primarily Vietnamese-news in-domain, so legal-domain reward use should not be interpreted as proof of legal-expert equivalence. This is why the paper complements it with a source-grounded GPT-5.5 judge and an independent human audit by lecturers teaching law courses.
Data used in the study
- Development/SFT/inner data:
duyet/vietnamese-legal-instruct, derived from Vietnamese legal documents collected from the national legal-document portal (vbpl.vn). The study retains the summarize task, which provides structured 3--5 sentence targets, and splits by source_id to avoid document-level leakage. The reported study uses 8,000 source--target pairs for SFT, 500 validation pairs, 5,000 source-only documents for CEFC-RFT reward selection, and 500 held-out inner-test documents. The reference targets are used for SFT/validation and reference-based metrics, but not for CEFC-RFT candidate scoring or pseudo-label selection.
- Distribution-shift test: 392 cleaned examples from Task 4.1 (legal-document summarization) of
VLegal-Bench. This benchmark is evaluation-only and its legal-document summarization references are expert-produced/cross-checked.
The current manuscript does not claim a literature SOTA on the duyet/vietnamese-legal-instruct inner split because we did not identify a published standardized summarization leaderboard for that exact task/split.
Evaluation
The final model was evaluated on:
- Inner test: 500 Vietnamese legal documents from the same general distribution.
- Out test: 392 documents under distribution shift.
Source-grounded GPT-5.5 evaluation
An anonymized source-summary pair was scored independently for Faithfulness (F), Coherence (C), and Relevance (R) on an integer 1–5 scale using a frozen Vietnamese legal-evaluation rubric. The reported judge configuration was ChatGPT GPT-5.5 in High reasoning mode.
Table with columns: Split, F, C, R, Overall| Split | F | C | R | Overall |
|---|
| Inner | 3.94 | 3.92 | 3.96 | 3.94 |
| Out | 3.60 | 4.34 | 3.08 | 3.68 |
Automatic metrics
ROUGE is reported as F1. BERTScore uses vinai/phobert-large after Vietnamese word segmentation.
Table with columns: Split, ROUGE-1 F1, ROUGE-2 F1, ROUGE-L F1, BERTScore F1| Split | ROUGE-1 F1 | ROUGE-2 F1 | ROUGE-L F1 | BERTScore F1 |
|---|
| Inner | 0.634 | 0.526 | 0.533 | 0.942 |
| Out | 0.171 | 0.137 | 0.148 | 0.903 |
The out-test references are much longer than the generated summaries, so low ROUGE recall should not be interpreted as factual failure by itself. Source-grounded evaluation and reference-based metrics should be considered jointly.
Reproducible source-grounded evaluation protocol
The paper does not use a reference summary for source-grounded legal evaluation. The full legal source and one anonymized model summary are evaluated together.
Frozen GPT-5.5 judge prompt
The following Vietnamese instruction is the frozen prompt used for the reported source-grounded LLM evaluation (ChatGPT GPT-5.5 Thinking, High reasoning mode, accessed September 2026):
Bạn là chuyên gia đánh giá văn bản pháp lý tiếng Việt và chất lượng đầu ra của mô hình ngôn ngữ.
Nhiệm vụ của bạn là đối chiếu ĐẦU RA MÔ HÌNH với VĂN BẢN PHÁP LÝ GỐC, sau đó đánh giá độc lập theo ba tiêu chí:
1. Trung thực (Faithfulness)
2. Mạch lạc (Coherence)
3. Liên quan (Relevance)
Mỗi tiêu chí phải được chấm bằng một số nguyên từ 1 đến 5. Không được sử dụng điểm thập phân.
Nguyên tắc chung
- Chỉ sử dụng thông tin có trong văn bản pháp lý gốc để đánh giá.
- Không sử dụng kiến thức bên ngoài để bổ sung, suy đoán hoặc sửa chữa nội dung.
- Không mặc định rằng đầu ra mô hình là đúng.
- Kiểm tra kỹ: chủ thể/đối tượng; quyền, nghĩa vụ, trách nhiệm; hành vi được phép/bắt buộc/bị cấm; điều kiện/ngoại lệ; phạm vi; thời hạn/hiệu lực; thẩm quyền; số liệu, mức tiền, tỷ lệ, ngày tháng, số điều/khoản/điểm; chế tài/hậu quả pháp lý; và quan hệ giữa các điều, khoản hoặc văn bản viện dẫn.
- Đặc biệt chú ý modality pháp lý như “phải”, “được”, “không được”, “có quyền”, “có trách nhiệm”, “chỉ khi”, “trừ trường hợp”, “trong thời hạn”, “tối đa”, “tối thiểu”, “có thể” và “theo quy định”.
- Ba tiêu chí phải được chấm độc lập.
Faithfulness (1–5)
Đánh giá mức độ mọi thông tin trong đầu ra được nguồn hỗ trợ chính xác. Kiểm tra hallucination/suy diễn, sai chủ thể hoặc thẩm quyền, đổi quyền thành nghĩa vụ, sai modality, mất điều kiện/ngoại lệ gây đổi nghĩa, sai số liệu/thời hạn/căn cứ, ghép quy định thành kết luận không được hỗ trợ, đảo ngược quy định, hoặc diễn đạt gây sai nghĩa pháp lý.
Việc lược bỏ chỉ làm giảm Faithfulness khi sự lược bỏ làm phần được trình bày trở nên sai hoặc gây hiểu nhầm; nếu phần còn lại vẫn đúng nhưng thiếu nội dung quan trọng, trừ ở Relevance.
Coherence (1–5)
Đánh giá tổ chức và diễn đạt của chính đầu ra: ngữ pháp, dễ hiểu, sắp xếp ý, liên kết, nhất quán chủ thể, tham chiếu rõ, không mâu thuẫn nội bộ hoặc lặp không cần thiết. Không trừ điểm chỉ vì đầu ra ngắn.
Relevance (1–5)
Đánh giá việc lựa chọn nội dung: mục đích và nội dung cốt lõi, phạm vi, chủ thể, quyền/nghĩa vụ, điều kiện, thủ tục/chế tài; phạt bỏ sót quan trọng, trọng tâm lệch hoặc chi tiết thứ yếu/thừa. Nội dung có thể liên quan nhưng vẫn không trung thực nếu trình bày sai.
Quy trình bắt buộc
1. Xác định nội dung pháp lý cốt lõi trong nguồn.
2. Đối chiếu từng claim kiểm chứng được của summary với nguồn.
3. Xác định lỗi thêm mới, sai lệch, suy diễn, mâu thuẫn hoặc bỏ sót.
4. Chấm F/C/R riêng.
5. Nhận xét ngắn, cụ thể, có căn cứ.
6. Nếu có lỗi, ghi claim sai/thiếu và source evidence.
7. Không xuất chain-of-thought; chỉ xuất kết quả cuối.
Chỉ trả về JSON hợp lệ:
{
"faithfulness_score": 1,
"faithfulness_comment": "...",
"coherence_score": 1,
"coherence_comment": "...",
"relevance_score": 1,
"relevance_comment": "...",
"major_errors": [
{
"model_claim": "...",
"source_evidence": "...",
"error_type": "hallucination | contradiction | wrong_subject | wrong_modality | wrong_condition | wrong_exception | wrong_number | wrong_time | wrong_authority | unsupported_inference | omission_causing_distortion | other"
}
],
"important_omissions": ["..."]
}
<VAN_BAN_PHAP_LY_GOC>
{{VĂN_BẢN_PHÁP_LÝ_GỐC}}
</VAN_BAN_PHAP_LY_GOC>
<DAU_RA_MO_HINH>
{{ĐẦU_RA_MÔ_HÌNH}}
</DAU_RA_MO_HINH>
Human-audit guide
The independent human audit follows the same source-only, criterion-separated design. The raters are university lecturers teaching law courses. They are not shown system identity, reference summaries, GPT-5.5 scores, or other raters' judgments.
Table with columns: Field, Human instruction| Field | Human instruction |
|---|
| Faithfulness (1–5) | Verify whether every presented legal claim is supported by the source. Omission reduces F only when it distorts the meaning of what is stated. |
| Coherence (1–5) | Judge clarity, grammaticality, organization, logical linkage, subject consistency, and absence of internal contradiction/repetition. |
| Relevance (1–5) | Judge coverage of legally important content and focus; missing important but non-distorting content lowers R. |
Major_error? | Mark YES if at least one legally material error is present; record Error_type, Model_claim, and Source_evidence. |
|
Required workflow: identify the core legal content → check each model claim against the source → record legal errors/omissions → score F/C/R independently → provide concise evidence-grounded comments.
Human audit results
The matched audit uses 98 sources (43 inner, 55 out), with the same source IDs for Base, SFT, RL-lite-500, and CEFC-RFT (392 source-summary pairs).
Table with columns: Split, Model, F, C, R, Overall, Major error, Important omission| Split | Model | F | C | R | Overall | Major error | Important omission |
|---|
| Inner | Base | 1.70 | 2.98 | 2.74 | 2.47 | 97.7% | 74.4% |
| Inner | SFT | 1.86 | 2.77 | 2.60 |
The human audit therefore supports a safety-informativeness trade-off rather than universal superiority: CEFC-RFT strongly improves inner relevance and reduces major-error incidence, while RL-lite-500 remains more coherent on the audited OOD subset.
Representative qualitative audit cases
These cases are included here rather than in the main paper to keep the conference manuscript within its page limit.
Success 1 — i206 (inner)
Topic: commercial-housing / social-housing land-allocation criteria.
- Base/SFT introduce unsupported legal identifiers.
- RL-lite-500 is safer but omits central provisions; expert score F/C/R = 3/4/2.
- CEFC-RFT preserves the key scope and land-allocation provisions without the unsupported metadata; expert score 5/5/4.
Interpretation: conditional feasibility-first selection can preserve source grounding without becoming as coverage-conservative as the pilot.
Success 2 — d344 (out)
Topic: policies for attracting and using talented personnel.
- Base, SFT, and RL-lite-500 introduce unsupported decree numbers and/or dates.
- CEFC-RFT preserves scope, target groups, implementation principles, and eligibility conditions; expert score 5/5/4.
Interpretation: the source-grounding mechanism can transfer to a distribution-shifted legal document even when baseline outputs remain fluent but metadata-unsafe.
Failure — d58 (out)
Topic: expenditure/funding rules for international treaties and agreements.
- CEFC-RFT captures the overall topic and important content but incorrectly associates several article numbers with the propositions assigned to them; expert score 2/3/4, with a major contradiction.
- Base and RL-lite-500 are more faithful on this particular instance.
Interpretation: literal metadata matching is not sufficient for article-proposition binding. A future verifier should jointly validate the cited article number and the proposition attributed to that article.
Recommended inference
Install:
pip install -U "transformers>=4.51.0" peft accelerate safetensors torch
Then load the base model and this adapter:
import re
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel
BASE_MODEL = "Qwen/Qwen3-0.6B"
ADAPTER_MODEL = "phuongntc/qwen3-0.6b-cefc-rft-vietnamese-legal"
dtype = (
torch.bfloat16
if torch.cuda.is_available() and torch.cuda.is_bf16_supported()
else torch.float16
if torch.cuda.is_available()
else torch.float32
)
tokenizer = AutoTokenizer.from_pretrained(ADAPTER_MODEL, use_fast=True)
base_model = AutoModelForCausalLM.from_pretrained(
BASE_MODEL,
torch_dtype=dtype,
device_map="auto" if torch.cuda.is_available() else None,
)
model = PeftModel.from_pretrained(base_model, ADAPTER_MODEL)
model.eval()
if tokenizer.pad_token_id is None:
tokenizer.pad_token = tokenizer.eos_token
Recommended source-grounded prompt:
SYSTEM_PROMPT = (
"Bạn là trợ lý tóm tắt văn bản pháp lý tiếng Việt. "
"Chỉ sử dụng thông tin có trong văn bản nguồn. "
"Ưu tiên tính trung thực, liên quan và chính xác pháp lý. "
"Không tự thêm số hiệu, ngày tháng, hiệu lực, cơ quan ban hành "
"hoặc kết luận pháp lý nếu văn bản nguồn không nêu rõ."
)
def build_messages(document):
return [
{"role": "system", "content": SYSTEM_PROMPT},
{
"role": "user",
"content": (
"Hãy tóm tắt ngắn gọn văn bản pháp lý sau. "
"Tập trung vào nội dung cốt lõi, chủ thể, phạm vi, "
"hành động pháp lý, nghĩa vụ/quyền hạn và điều kiện "
"quan trọng nếu có.\n\n"
f"VĂN BẢN NGUỒN:\n{document}\n\n"
"TÓM TẮT:"
),
},
]
For deterministic inference, use greedy decoding (do_sample=False). A complete script is provided as inference_cefc_rft.py.
Important implementation note
For decoder-only generation, decode only tokens generated after the full input width:
prompt_width = inputs["input_ids"].shape[1]
generated_ids = outputs[0][prompt_width:]
summary = tokenizer.decode(generated_ids, skip_special_tokens=True)
Do not slice using attention_mask.sum() when left padding is used, because that can leak prompt/source tokens into the decoded output.
Intended use
This model is intended for:
- research on Vietnamese legal summarization;
- source-grounded legal-document compression;
- experiments on factuality, legal metadata grounding, and offline reward-guided fine-tuning;
- academic comparison with Base, SFT, and other reward-guided systems.
Out-of-scope use
This model should not be used as:
- a substitute for a lawyer, regulator, or authorized legal professional;
- an autonomous legal decision-making system;
- a source of binding legal interpretation;
- a tool for producing legal advice without verification against the original document.
Limitations
The model has several known limitations:
- It was evaluated on Vietnamese legal documents using one compact backbone.
- The reported CEFC-RFT run uses one random seed.
- A source-grounded LLM judge can still make evaluation errors and is not a substitute for expert legal review.
- Long documents may contain important information outside the model's practical prompt budget.
- The out-of-distribution test shows that relevance/coverage remains more difficult than coherence and faithfulness.
- Concise summaries may omit legally important secondary provisions even when the generated statements are correct.
- Deterministic metadata checks reduce but do not eliminate unsupported legal identifiers, dates, authorities, or effect statements. A documented failure shows that a correct article number may still be attached to an incorrect proposition, motivating article-proposition binding verification.
- EViSE-mmBERT is a proxy reward model whose published validation is primarily Vietnamese-news in-domain; legal-domain use is independently audited but is not equivalent to expert legal certification.
Always verify a generated summary against the original legal source before operational use.
Training details
The final CEFC-RFT continuation uses margin-weighted assistant-only maximum likelihood on selected pseudo-labels. Source and prompt tokens are masked from the training loss.
Exported metadata:
method: CEFC-RFT margin-weighted assistant-only fine-tuning
training examples: 2,394
CEFC pseudo-labels: 2,394
SFT anchor examples: 0
epochs: 1
learning rate: 2e-5
gradient accumulation: 16
The final reward-selection run used:
N_SELECT = 5000
w_R = 0.30
w_C = 0.70
epsilon = 0.36821579933166504
metadata feasible threshold = 0.25
lambda_0 = 1.5
mu_metadata = 0.5
delta = 0.03
eta_lambda = 0.05
final lambda = 1.512151161492657
stage-1 relevance gate = 0.28
stage-1 coherence gate = 0.48
stage-2 relevance gate = 0.32
stage-2 quality gate = 0.42
max new tokens during reward-selection generation = 180
License
The adapter is released under the Apache License 2.0, consistent with the base model license. Users are responsible for complying with the licenses and access conditions of any data used in downstream applications.