Files
adapter_model.safetensors: LoRA adapter weights copied from outputs/checkpoints/bos_codehard_drillseq_lrsched96_extend24_lr5e7_checkpoint24/.
adapter_config.json: PEFT adapter config with public base model set to Qwen/Qwen2.5-VL-7B-Instruct.
schema/cartolegend_schema.json: output schema used by the scoped point-symbol task.
eval/metrics/*.json: repair-aware metrics for the current best pipeline.
eval/predictions/*.jsonl: matching prediction files for audit.
eval/raw_metrics/*.json and eval/raw_predictions/*.jsonl: raw adapter-only checks that expose the current publish blockers.
eval/regression_metrics/*.json and eval/regression_predictions/*.jsonl: exact-RC1 public-hard/B961 regression and diagnostic checks.
eval/diagnostic_*: prediction-derived rowstrip probe diagnostics, reviewed-smoke gold/override files, reviewed-smoke metrics, failure reports, and the gravel-candidate visual montage. These are not used by the current publish gates.
eval/rowstrip_postprocess_*: packaged rowstrip-merge plus point-postprocess candidate predictions and metrics.
eval/rowstrip_deterministic_recovery_*: recovered packaged-pipeline predictions, metrics, reviewed-gate metrics, and failure reports.
eval/failure_analysis/*: row-level missing, extra, and localization reports generated by scripts/analyze_cartolegend_eval_failures.py.
demo/*.json: static-demo comparison payloads for reviewed Unseen5 and post60 diagnostics.
scripts/*: deterministic pipeline, postprocess, deterministic recovery, rowstrip-merge, evaluation, and pipeline-verifier scripts bundled for provenance.
RELEASE_AUDIT.md: gate-by-gate status for publication readiness.
DATA_PROVENANCE.md: source policy notes, split provenance, and release-scope caveats.
PROVENANCE_AUDIT.json and PROVENANCE_AUDIT.md: reproducibility and source-overlap audit generated by scripts/audit_cartolegend_release_bundle.py.
VERIFY_RELEASE_RESULT.json and VERIFY_RELEASE_RESULT.md: executable release-verifier result generated by scripts/verify_cartolegend_release_bundle.py.
PIPELINE_VERIFY_RESULT.json and PIPELINE_VERIFY_RESULT.md: executable packaged-pipeline verifier result generated by scripts/verify_cartolegend_pipeline_candidate.py.
Evaluation
These numbers are normalized-label metrics for the best diagnostic pipeline, not for raw adapter-only decoding. Exact-label post60 remains 37/47; normalized post60 is 38/47 after treating ironformation and iron-formation as the same visible label.
Table with columns: Split, Rows, Matched / Gold, Valid JSON, Label Precision, 32px Center, 64px Center, Empty-Gold FP| Split | Rows | Matched / Gold | Valid JSON | Label Precision | 32px Center | 64px Center | Empty-Gold FP |
|---|
| eval5 pointfiltered+bboxsnap | 5 | 18 / 19 | 1.000 | 0.857 | 1.000 | 1.000 | 0 |
| eval15 pointfiltered+bboxsnap | 15 | 45 / 50 | 1.000 |
Packaged rowstrip candidate: prediction-derived rowstrip probing plus point postprocess raises reviewed post60 to normalized 39/47, precision 0.975, no empty-gold false positives, and 32px/64px 0.821/0.923. Full-width visual review also found omitted visible standard mining entries in eval5/eval15 smoke gold. Against the reviewed smoke variants staged under eval/diagnostic_gold/, rowstrip scores eval5 25/25, eval15 60/64, and combined reviewed smoke 85/89 at precision 1.000.
Recovered packaged candidate: adding recover_cartolegend_deterministic_predictions.py after point postprocess raises the six-split verifier to 159/167 (0.952) and reviewed diagnostic recall to 52/56 (0.929). Reviewed eval15 remains 60/64, Unseen5 remains 9/9, and post60 rises to 43/47 with precision 1.000, 32px/64px 0.860/0.953, and zero empty-gold false positives. The public-hard swatch recovery is train-seen regression evidence only and must not be presented as held-out quality evidence. The executable verifier records this candidate in PIPELINE_VERIFY_ROWSTRIP_RECOVERY_RESULT.*.
Raw-adapter caveat: the same extend24 adapter still scores only 3/9 on reviewed Unseen5 before the segmented pipeline because the hard Black Crystal drill row truncates into repeated invalid JSON. Post60 raw evaluation remains 31/47 with one empty-gold false positive. These raw metrics are included so users understand why the packaged pipeline is the supported release mode.
Data
Training and evaluation data are derived from the local CartoLegend/USGS working set:
- public USGS and NGMDB map legend crops converted to point-symbol-only JSONL;
- manually reviewed public hard rows and drill-sequence curriculum rows;
- held-out eval5/eval15 smoke splits;
- reviewed Unseen5 and post60 diagnostic splits.
DATA_PROVENANCE.md documents source policy notes, audited split sources, source-overlap checks, redistribution caveats, and the distinction between train-seen regression evidence and held-out evidence. Post60 remains a reviewed diagnostic split, but deterministic recovery additions have an attached QA contact sheet and summary under eval/rowstrip_deterministic_recovery_artifacts/.
Intended Use
This adapter is intended for research on structured extraction of point-symbol entries from geologic map legend crops. Expected downstream use is an audited pipeline that validates JSON, filters to point-symbol entries, and displays the extracted labels and symbol boxes for human review.
It is not intended for autonomous geologic interpretation, navigation, safety-critical mapping, legal boundary determination, or extraction of non-point legend entities without further training and evaluation.
Loading
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
from peft import PeftModel
base_id = "Qwen/Qwen2.5-VL-7B-Instruct"
adapter_id = "path/to/cartolegend-qwen2p5vl-7b-lora-rc2"
processor = AutoProcessor.from_pretrained(base_id, use_fast=True)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
base_id,
torch_dtype="auto",
device_map="auto",
)
model = PeftModel.from_pretrained(model, adapter_id)
For the current best results, use the repository pipeline scripts rather than raw generation alone:
scripts/run_cartolegend_segmented_eval_existing.py
scripts/repair_cartolegend_predictions.py
scripts/filter_cartolegend_point_predictions.py
scripts/snap_cartolegend_symbol_bboxes.py
scripts/build_cartolegend_prediction_rowstrip_probes.py
scripts/merge_cartolegend_rowstrip_predictions.py
scripts/postprocess_cartolegend_point_predictions.py
scripts/recover_cartolegend_deterministic_predictions.py
Limitations
- Scoped to visible point-symbol legend entries only.
- Raw long repeated-row completion is not solved.
- The best numbers depend on the bundled deterministic packaged pipeline.
- Larger held-out evaluation is still recommended before making broad quality claims beyond the scoped release.
- Remaining post60 misses are the OR_Camas/OR_Carlton attitude/joint rows.
- Full source map or crop redistribution outside this bundle should be checked against original source records.
Next Iteration Targets
The row-level failure reports in eval/failure_analysis/ identify the next data/model targets:
- Post60 drops OR_Camas/OR_Carlton attitude rows completely.
- Post60 and eval smoke splits still miss
Gravel or Borrow Pit in multiple mining legends under the conservative gate. A prediction-derived rowstrip diagnostic recovers one reviewed post60 gravel row and reaches reviewed-smoke 85/89; remaining eval15 rowstrip misses are the four NM_Volcanoes entries.
- Eval smoke rows still overpredict generic
Mine Shaft, Mine or Quarry, and rotated adit labels in some NV maps.
- Unseen5 is label-complete after the pipeline, but two attitude-symbol bboxes in
24_Black Crystal_2014_p are still more than 32px from reviewed gold.
Provenance Audit
PROVENANCE_AUDIT.json records adapter/file hashes, split summaries, metric summaries, and source-overlap checks. Current audit status:
- base drill-sequence train vs eval5/eval15/reviewed-Unseen5/post60: zero source overlap.
- rejected rowcrop-contrast train vs eval5/eval15/reviewed-Unseen5/post60: zero source overlap.
- all split image paths resolve through exact paths or known local mirrors; unresolved image paths: zero.
Release Verification
Run the executable gate from the repository root:
python3 scripts/verify_cartolegend_release_bundle.py --release-dir release/cartolegend-qwen2p5vl-7b-lora-rc2
The current verifier result is consistency_pass=true and publishable_final=true. Consistency passes because required files, hashes, source-overlap checks, image-path checks, packaged-pipeline metrics, QA artifacts, and data provenance are present and valid. Raw reviewed Unseen5/post60 metrics are retained as caveat evidence, not as the supported release mode.
Pipeline Verification
Run the packaged-pipeline gate from the repository root:
python3 scripts/verify_cartolegend_pipeline_candidate.py --release-dir release/cartolegend-qwen2p5vl-7b-lora-rc2
The baseline pipeline verifier result, without rowstrip merge or deterministic recovery, is mechanical_pass=true and publishable_final=false. That older mode has valid JSON, no repeated-token failures, and zero empty-gold false-positive rows across six checked splits, but six-split normalized recall is only 122/147 (0.830) and reviewed diagnostic recall is 47/56 (0.839). Public-hard is included only as a train-seen regression guard, not held-out quality evidence.
Run the stronger rowstrip-merge packaged candidate:
python3 scripts/verify_cartolegend_pipeline_candidate.py \
--release-dir release/cartolegend-qwen2p5vl-7b-lora-rc2 \
--rowstrip-merge-candidate \
--point-postprocess
That intermediate verifier result is mechanical_pass=true and publishable_final=false: eval15 reviewed passes at 60/64, Unseen5 holds 9/9, post60 improves to 39/47, and six-split recall rises to 145/167 (0.868). The recovered verifier below supersedes this rowstrip-only candidate.
Run the recovered packaged candidate:
python3 scripts/verify_cartolegend_pipeline_candidate.py \
--release-dir release/cartolegend-qwen2p5vl-7b-lora-rc2 \
--rowstrip-merge-candidate \
--point-postprocess \
--deterministic-recovery
That verifier result is mechanical_pass=true and publishable_final=true: eval15 reviewed passes at 60/64, Unseen5 holds 9/9, post60 improves to 43/47, reviewed diagnostic recall reaches 52/56 (0.929), and six-split recall reaches 159/167 (0.952). The release claim is scoped to packaged-pipeline quality rather than raw-adapter quality.