Summary
The main goal of this work was to improve the ARC-Challenge score of
Mistral-7B while preserving a reproducible and computationally efficient
training pipeline.
The final clean model improved ARC-Challenge test acc_norm from
0.6152 to 0.7662, corresponding to a 15.10 percentage-point gain.
Method
For each ARC question, every answer candidate is independently appended to
the evaluation prompt:
Question: {question}
Answer: {candidate}
The candidate score is defined as:
score(candidate) =
sum of candidate-token log probabilities / len(candidate_text)
Cross-entropy loss is then applied over all candidate scores so that the
correct answer is ranked above the incorrect answers.
This directly matches the character-length normalization used by the
acc_norm implementation in the evaluation harness.
Training Data
- Dataset:
allenai/ai2_arc
- Configuration:
ARC-Challenge
- Original training split: 1,119 questions
- Final clean training split: 1,112 questions
- Validation split: 299 questions
Questions with normalized exact-text overlap with the ARC-Challenge
validation or test split were removed from training.
Normalization consisted of:
- lowercasing
- whitespace normalization
- exact question-text matching
No validation or test answers were used for training.
Training Configuration
Table with columns: Setting, Value| Setting | Value |
|---|
| Base model | mistralai/Mistral-7B-v0.1 |
| Fine-tuning method | 4-bit NF4 QLoRA |
| Trainable modules | q_proj, v_proj |
| LoRA rank | 8 |
| LoRA alpha | 16 |
| LoRA dropout | 0.05 |
| Bias | None |
| Epochs |
Evaluation Protocol
Evaluation was performed using EleutherAI's
lm-evaluation-harness.
- Task:
arc_challenge
- Number of few-shot demonstrations: 25
- Evaluation dtype: FP16
- Batch size: 8
- Tokenizer: slow tokenizer (
use_fast_tokenizer=False)
- Harness commit:
ae79b1217aad7738b91e88a4017c86a5d5e45aa7
The final method was selected using the ARC-Challenge validation split.
The test split was evaluated after the method and hyperparameters had been
fixed.
Main Results
Table with columns: Model, Validation acc, Validation acc_norm, Test acc, Test acc_norm| Model | Validation acc | Validation acc_norm | Test acc | Test acc_norm |
|---|
| Mistral-7B base | 0.5184 | 0.5786 | 0.5700 | 0.6152 |
| Final clean ranking adapter | 0.7023 | 0.7391 | 0.7415 | 0.7662 |
Improvement over the base model
Table with columns: Metric, Absolute improvement| Metric | Absolute improvement |
|---|
| Validation acc | +18.39 percentage points |
| Validation acc_norm | +16.05 percentage points |
| Test acc | +17.15 percentage points |
| Test acc_norm | +15.10 percentage points |
Ablation Summary
Table with columns: Experiment, Main change, Validation acc, Validation acc_norm| Experiment | Main change | Validation acc | Validation acc_norm |
|---|
| Base | No fine-tuning | 0.5184 | 0.5786 |
| Completion-only SFT | Correct-answer language-modeling loss | 0.5485 | 0.5786 |
| Ranking, seed 42 | Metric-aligned ranking loss | 0.7157 | 0.7525 |
| Ranking, seed 123 | Same method, different seed | 0.6890 | 0.7224 |
Across seeds 42, 123, and 427, the ranking objective achieved:
Validation acc_norm: 0.7391 ± 0.0153
Validation acc: 0.6990 ± 0.0146
The reported standard deviation is the sample standard deviation over the
three seeds.
Loading the Adapter
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base_model_id = "mistralai/Mistral-7B-v0.1"
adapter_model_id = "yunagreenpark/Mistral-7B-v0.1-ARC"
tokenizer = AutoTokenizer.from_pretrained(
adapter_model_id,
use_fast=False,
)
if tokenizer.pad_token_id is None:
tokenizer.pad_token = tokenizer.eos_token
base_model = AutoModelForCausalLM.from_pretrained(
base_model_id,
torch_dtype=torch.float16,
device_map="auto",
)
model = PeftModel.from_pretrained(
base_model,
adapter_model_id,
)
model.eval()
Evaluation Example
After installing the same lm-evaluation-harness version, the adapter can
be evaluated as follows:
lm_eval \
--model hf \
--model_args "pretrained=mistralai/Mistral-7B-v0.1,peft=yunagreenpark/Mistral-7B-v0.1-ARC,trust_remote_code=True,dtype=float16,use_fast_tokenizer=False" \
--tasks arc_challenge \
--device cuda:0 \
--batch_size 8 \
--num_fewshot 25
Repository Files
adapter_model.safetensors: trained LoRA adapter weights
adapter_config.json: PEFT adapter configuration
tokenizer.model: SentencePiece tokenizer model
tokenizer_config.json: tokenizer configuration
special_tokens_map.json: special-token mapping
training_args.bin: serialized training arguments