Model Description
This model is a full-parameter fine-tuned version of
HuggingFaceTB/SmolLM-135M trained on chemistry
SMILES strings from the
Codemaster67/Causal_lm_chemistry_1M_rows dataset.
The base model's tokenizer was pre-extended with ~300 SPE (SMILES Pair
Encoding) chemistry tokens plus <|start_of_smiles|> / <|end_of_smiles|>
special tokens, and its embedding & LM-head layers were resized with
mean-initialised vectors for the new tokens.
Training Details
Table with columns: Parameter, Value| Parameter | Value |
|---|
| Method | Full Fine-Tune (all weights updated) |
| Parallelism | FSDP (Fully Sharded Data Parallel) |
| Epochs | 1 |
| Learning Rate | 5e-06 |
| Batch Size (per device) | 32 |
| Gradient Accumulation | 1 |
| Max Sequence Length | 128 |
| Warmup Ratio | 0.1 |
| Weight Decay | 0.01 |
| Scheduler | Cosine |
| Precision | bf16 |
| Augmentation | OFF |
| Training Samples | Full dataset |
| Eval Samples | Full dataset (10%) |
Evaluation Results
Table with columns: Metric, Value| Metric | Value |
|---|
| Final Eval Loss | 3.204535961151123 |
| Final Eval Perplexity | 24.644061560759287 |
| Training Loss | 3.1543 |
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Codemaster67/Test_run", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("Codemaster67/Test_run", trust_remote_code=True)
smiles_input = "<|start_of_smiles|>CC(=O)Oc1ccccc1C(=O)O<|end_of_smiles|>"
inputs = tokenizer(smiles_input, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))
Intended Use
Chemistry-domain language modelling, SMILES generation and completion,
and downstream molecular property prediction via fine-tuning.
Limitations
- Trained primarily on SMILES strings; natural-language instruction-following
ability may degrade compared to the base OLMo checkpoint.
- Augmentation was disabled for this run.