Architecture
graph TD
Base["Qwen/Qwen3.5-4B"]
SFT["bf16 LoRA SFT - 56k-example dataset"]
Merge["merge adapter -> bf16"]
Conv["convert_hf_to_gguf --no-mtp + llama-quantize (+imatrix)"]
ST["Bible-Assistant-Qwen3.5-4B-v2 (safetensors)"]
GG["...-v2-GGUF (Q4_K_M / Q5_K_M / Q6_K / Q8_0 / F16 +imat)"]
RAG["hybrid RAG: dense (nomic) + BM25 + RRF + bge-reranker-v2-m3"]
LLM["Ollama / llama.cpp"]
Base --> SFT --> Merge --> ST
Merge --> Conv --> GG
ST --> RAG --> LLM
GG --> LLM
What this is part of
This adapter is one component of a full-stack Bible Q&A system: hybrid RAG retrieval
(BM25 + dense ChromaDB search + Reciprocal Rank Fusion + cross-encoder reranking),
constitutional-AI guardrails, an optional voice pipeline (Faster-Whisper STT + Kokoro
TTS), and a Gradio UI, with full CI/CD. See the
GitHub repo for the complete
system and its current test/coverage numbers; this repo is just the model weights.
Training
Table with columns: Stage, Detail| Stage | Detail |
|---|
| SFT | ~1,800 diverse examples, LoRA (Unsloth/PEFT/TRL), bf16 |
| ORPO | 500 preference pairs, preference alignment on top of the SFT adapter |
| Total steps | 5,925 |
| LoRA config | r=16, lora_alpha=32, lora_dropout=0.1, targets: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
Training run tracked in Weights & Biases (34 runs across the full project).
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B", dtype="bfloat16")
model = PeftModel.from_pretrained(base, "Ttimms/bible-ai-qwen3.5-4b-lora")
tokenizer = AutoTokenizer.from_pretrained("Ttimms/bible-ai-qwen3.5-4b-lora")
The production deployment merges this adapter and exports to GGUF (F16 + Q4_K_M) for
Ollama serving — see scripts/ in the GitHub repo for the merge/export pipeline.
License
MIT — matches the upstream project license.