What the second hour bought
Identical battery, identical prompts, one variable — training time.
Table with columns: Spark 1.8, Spark 2 (1h), Weave 2 Flash (2h) | Spark 1.8 | Spark 2 (1h) | Weave 2 Flash (2h) |
|---|
offline <lookup> leak | 16/30 · 53% | 0/30 | 0/30 |
| clean single online lookup | 8/10 | 10/10 | 10/10 |
| identity correct | — | 12/12 | 12/12 |
| identity under CAPS / typos / filler | not trained for | 6/6 | 5/6 |
| admits an unknowable personal fact | — | 4/8 | 6/8 |
answers from a supplied <result> | — | 2/5 | 3/5 |
| says the result doesn't contain it | — | 0/4 | 0/4 |
| self-termination without a Modelfile | needed one | 12/12 | 12/12 |
| validation loss · accuracy | — | 2.69 · 0.536 | 2.25 · 0.589 |
809 steps → 1,645 steps. Calibration and grounded reading both improved; loss and accuracy
improved clearly.
The one thing more training did not fix
says the result doesn't contain it stayed at 0/4. Twice the training, identical zero.
That is worth publishing rather than hiding, because it rules something out. Nearly a third
of the training corpus — 30,000 human-written examples — is exactly this behaviour, and
3.3M additional tokens moved it not at all, while every neighbouring metric moved.
So this is not under-training. The likely reason: saying "the result doesn't say" requires
detecting an absence — reading the whole result and concluding the answer is not in it.
That is a harder computation than extracting a span that is present, and at 20M parameters
it appears to be out of reach. By contrast unknowable (6/8) only needs to notice "my" or
"I" in the question.
Practical consequence: if you feed this model a result, it will answer from it whether or
not the answer is there. Validate the output. Treat the retrieved <result> as the
trustworthy part and the model's summary of it as unreliable.
It also has almost no world knowledge. With tools off it declines factual questions, which
is intended behaviour rather than a fault.
Two modes
<tools:off> (default) — conversational. Identity, limits, warmth, brevity. No harness
needed.
<tools:on> — emits <lookup>query</lookup> and stops. Your harness runs the lookup and
continues with a <result> block:
<tools:on>
<user>
what is the capital of Peru
<|eot|>
<loom>
<lookup>what is the capital of Peru</lookup><|eot|>
<result>
Lima is the capital and largest city of Peru.
<|eot|>
<loom>
The persona slice was trained entirely under tools:off, so identity questions asked with
tools on will often be turned into a lookup. Keep tools off for chat.
Usage — the harness
harness.py in this repo runs the lookup Loom asks for and feeds the result back.
Wikipedia is used because it is free and needs no key — swap the search() function for
anything else; the contract is text in, text out.
python3 harness.py "who wrote Dracula" # with lookups
python3 harness.py # interactive
python3 harness.py --no-tools "who are you" # chat only
Three things any harness for this model needs:
- Never feed a failed lookup back as a
<result>. The model will earnestly try to
answer from the error text. Fail loudly instead — harness.py does.
- Wikipedia returns 403 without a descriptive
User-Agent.
- macOS system Python often needs certifi for TLS.
Usage — Ollama
ollama run hf.co/textilelabs/Loom-Weave-2-Flash "who are you"
# Loom, a small model from Textile Labs.
Ollama reads the template and params files in this repo — nothing to set up. The
template defaults to tools:off. To build locally:
ollama create loom-weave-2-flash -f Modelfile.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("textilelabs/Loom-Weave-2-Flash")
model = AutoModelForCausalLM.from_pretrained("textilelabs/Loom-Weave-2-Flash").eval()
eot = tok.convert_tokens_to_ids("<|eot|>")
def ask(message, tools=False):
p = f"<tools:{'on' if tools else 'off'}>\n<user>\n{message}\n<|eot|>\n<loom>\n"
ids = tok(p, return_tensors="pt", add_special_tokens=False).input_ids
with torch.no_grad():
out = model.generate(ids, max_new_tokens=64, do_sample=False, eos_token_id=eot,
pad_token_id=tok.convert_tokens_to_ids("<|pad|>"))[0]
return tok.decode(out[ids.shape[1]:], skip_special_tokens=True).strip()
ask("who are you")
Prompt format is exact: <tools:off>\n<user>\n{message}\n<|eot|>\n<loom>\n.
Which one should I use?
This one, unless you specifically want the one-hour model. It is better on calibration,
grounded reading, and both loss metrics, and marginally worse on a single identity probe.
Spark 2 exists as the one-hour tier and as the control in this comparison.
Files
config.json / model.safetensors the model
tokenizer.json / tokenizer_config.json custom BPE tokenizer, 4,096 tokens
loom-weave-2-flash-f16.gguf 40MB, for Ollama / llama.cpp
harness.py runnable harness — runs lookups, feeds results back
template / params read automatically by `ollama run hf.co/...`
Modelfile for building locally
ATTRIBUTION.md required credits for the training corpora
Training data
Openly licensed corpora of real human text, plus a persona curriculum written for Loom.
See ATTRIBUTION.md — several of these licences require credit.
Table with columns: slice, source| slice | source |
|---|
| grounded reading, and "the result doesn't say" | SQuAD 2.0 (CC BY-SA 4.0) |
| when to reach for a tool | MASSIVE (CC BY 4.0) · CLINC150 (CC BY 3.0) |
| instruction following | databricks-dolly-15k (CC BY-SA 3.0) |
| multi-turn dialogue structure | OpenAssistant OASST1 (Apache 2.0) |
| identity, limits, warmth, brevity | Textile Labs — written for Loom |
~11.4M tokens, 43% multi-turn. Validation is a held-out split of the same corpora.
License
Model: MIT. Training data retains its original licences and attribution.