Introduction
Lythri is a family of on-device language models built for emotional companionship. Instead of chasing math and coding scores, Lythri is trained to understand how people feel and to hold natural, multi-turn conversations, while staying small enough to run locally on a laptop or phone.
Emotional Intelligence
Zero-shot comparison between Lythri-4B-A2B and Gemma 4 instruction-tuned baselines on three emotion benchmarks. Despite having only 2.3B active parameters, Lythri-4B-A2B(4.6B, base on Gemma4-E2B-PT) comparable to Gemma4-E4B-IT(8B) on all three EQ benchmarks.
Note: For full details including Lythri-7B-A4B, please refer to the technical report.
General Benchmarks
All benchmarks are evaluated with their official standard settings and in generative mode with chat template applied, reflecting real-world inference conditions. Think-tag outputs from model are stripped before answer extraction.
Quickstart
from transformers import AutoModelForCausalLM, AutoTokenizerimport torch model_path = "Lythri/Lythri-7B-A4B" # or "Lythri/Lythri-4B-A2B"tok = AutoTokenizer.from_pretrained(model_path)model = AutoModelForCausalLM.from_pretrained( model_path, dtype=torch.bfloat16, device_map="auto") messages = [{"role": "user", "content": "My friend just lost their job and seems really down. What should I say to them?"}]chat = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) + "<think>"inputs = tok(chat, return_tensors="pt").to(model.device) with torch.inference_mode(): out = model.generate( **inputs, max_new_tokens=2048, do_sample=False, eos_token_id=[1, 106], ) print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Recommended Generation Config
generation_config = { "temperature": 0.7, "top_p": 0.9, "top_k": 64, "max_new_tokens": 2048, "repetition_penalty": 1.05, "do_sample": True, "eos_token_id": [1, 106],} out = model.generate(**inputs, **generation_config)
System Prompt (Optional)
For best results in emotional support scenarios, we recommend:
You are Lythri, an AI assistant developed by Spike8086 for emotional support. Reply concisely in the same language as the user. Always consider the user's feelings. Be friendly and warm, and provide concrete help in critical moments.
The model works fine but not the best without a system prompt.
Compute
The full development of Lythri, including training and evaluation, used about 2,842 GPU hours on NVIDIA RTX 6000D GPUs.
Training on various open-source datasets, 20B data for CPT stage, 77K pairs for SFT, and 5000x3 responses for GRPO judge by Gemma2-27B.
Limitations
- Lythri is optimized for conversation and emotional understanding, not for math, coding or complex reasoning.
- Lythri is not a substitute for professional mental health support. If you or someone you know is in crisis, please contact local emergency services or a crisis helpline.
- Like all language models, it can produce inaccurate or inappropriate content.
License
Lythri is built on Gemma 4 and is released under the Apache License 2.0.
Citation
@techreport{li2026lythri, title = {Lythri Technical Report}, author = {Li, Jiawen}, year = {2026}, institution = {Zenodo}, doi = {10.5281/zenodo.23179311}, url = {https://doi.org/10.5281/zenodo.23179311}}
Support
The whole training process is self-funded. If you like our model, please click a free like — that means a lot to me as a student!