Why it exists
General instruction models are strong writers but unreliable at structured
output: a few percent of the time they wrap JSON in fences, leak reasoning, or
return the wrong shape/count — which breaks any automated pipeline. This model
closes that reliability gap for the marketing-content-plan task.
Results (100 held-out scenarios, judge-independent hard metrics)
Table with columns: Metric, Baseline, Qwen-Marketing-S1| Metric | Baseline | Qwen-Marketing-S1 |
|---|
| Aggregate score | 0.954 | 0.999 |
| Parses as JSON | 0.95 | 1.00 |
| Correct array shape | 0.95 | 1.00 |
| Post count 6–8 | 0.93 | 1.00 |
| Posts with all keys | 0.94 | 1.00 |
| Valid platforms | 0.95 | 1.00 |
| Valid content types | 0.95 | 1.00 |
| Hashtags 5–15 | 0.94 | 0.99 |
| Caption within platform limit | 0.94 | 1.00 |
Every gate reaches 100%: the outputs the baseline broke (invalid JSON on ~5%,
wrong post count on ~7%) drop to zero.
Output schema
Each element of the returned array:
{ "platform": "instagram", "content_type": "carousel", "caption": "...", "hashtags": ["...", "..."], "media_prompt": "a prompt for an image model", "reasoning": "why this post fits the brief"}
Valid platform: instagram, twitter/x, linkedin, facebook, tiktok.
Valid content_type: text, image, video, carousel, reel.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizerimport torch model_id = "AbdulrahmanOmar/qwen-marketing-s1"tokenizer = AutoTokenizer.from_pretrained(model_id)model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype=torch.bfloat16, device_map="auto") messages = [ {"role": "system", "content": "You are a senior content creator within an AI marketing platform. Output the deliverable immediately as one valid JSON array and nothing else."}, {"role": "user", "content": "Create a social media content calendar for a specialty coffee roaster launching a summer single-origin Ethiopian bean. Platforms: instagram, tiktok. Create 6-8 posts."},]inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors="pt").to(model.device)out = model.generate(inputs, max_new_tokens=2048, do_sample=False)print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
Trained with enable_thinking=False; keep thinking mode off at inference.
How it was trained
- Knowledge distillation → QLoRA SFT → merged into a standalone model.
- Teacher: a 14B instruction model (
Qwen/Qwen2.5-14B-Instruct-AWQ) generated
content plans over 1,400 synthetic brand/campaign scenarios; only schema-valid
outputs were kept as training targets (1,321 pairs; 94.4% valid).
- Fine-tuning: QLoRA (4-bit NF4 base), LoRA on all attention + MLP
projections, assistant-only loss, then merged to fp16 for release.
Table with columns: Setting, Value| Setting | Value |
|---|
| LoRA rank / alpha / dropout | 16 / 32 / 0.05 |
| Epochs | 3 |
| Learning rate / schedule | 2e-4 / cosine, 3% warmup |
| Max sequence length | 2560 |
| Effective batch size | 16 |
| Optimizer | paged AdamW 8-bit |
| Training examples | 1,321 |
Limitations
- Purpose-built for the JSON content-plan schema above — not a general chat model.
- English marketing scenarios only; evaluated in-distribution on synthetic data.
- Can still produce off-brand copy or unverified claims — keep a human in the loop.
License
Apache-2.0.
Base model
Fine-tuned from marketeam/Qwen-Marketing (Apache-2.0).