Skip to content

ADR-814: SigLIP 2 as the default local embedding profile

Context

The default embedding profile (Nomic v1.5, seeded by migration 056 and carried in schema/00_baseline.sql) loads nomic-ai/nomic-embed-text-v1.5 and nomic-ai/nomic-embed-vision-v1.5 with trust_remote_code. Both models run on the nomic-ai/nomic-bert-2048 remote code, last changed 2025-04-29, and that code is broken on transformers 5.x (issue #565, PR #552):

  • the vision model's config.json carries "n_inner": 2048.0, which the 5.x strict config dataclass rejects;
  • the text model calls get_extended_attention_mask, removed from PreTrainedModel in 5.x.

requirements.txt therefore caps transformers<5.0, and the cap transitively holds sentence-transformers at 5.x. The cap is the last held dependency from the 2026-09-17 sweep. Waiting on an upstream fix has no date attached.

ADR-803 settled the shape of the embedding system: one universal text space, and a modality's native embedding is an independent same-modality index that is never compared to the text space. That decoupling means the text model and the image model can be replaced independently, and neither needs to be a single co-spatial multimodal model.

Verified on 2026-09-18 with transformers 5.17.0, sentence-transformers 6.1.0, torch 2.14.0 (CPU):

Model Role Dims Context Remote code License Weights
nomic-ai/modernbert-embed-base text 768 8192 tokens no (model_type: modernbert) Apache 2.0 596 MB fp32
google/siglip2-base-patch16-256 image 768 256 px no (model_type: siglip) Apache 2.0 1.5 GB fp32 (both towers)

ModernBERT-embed is Nomic's own successor to nomic-embed-text: same 768 dimensions, same search_query: / search_document: task prefixes, same 8K context, native transformers architecture. The SigLIP 2 vision tower's pooler_output matches the full model's get_image_features exactly, so the image index only needs the vision tower loaded.

SigLIP 2's text tower is limited to 64 tokens. It is not usable as the universal text embedder for 1000-word source chunks, which rules out the single-model "multimodal" profile shape that the earlier plan (commit 3649c73a5) sketched.

Decision

  1. The default local embedding profile becomes ModernBERT + SigLIP 2 (migration 081): text slot nomic-ai/modernbert-embed-base via sentence-transformers, image slot google/siglip2-base-patch16-256 via transformers, multimodal = false, trust_remote_code = false on both slots, vector space modernbert-embed-base, image vector space siglip2-base-p16-256.

  2. The transformers<5.0 cap is lifted. requirements.txt floors transformers>=5.0 and sentence-transformers>=6.0.

  3. Migration 081 switches existing installs whose active profile is the seeded Nomic v1.5 row to the new profile, mirroring the activate_embedding_config flow (protection flags, vocabulary embeddings marked stale). The Nomic row stays in the table, inactive and unprotected. Installs on any other active profile (OpenAI, custom) are left alone.

  4. Existing text embeddings are re-generated by the operator, not by the migration:

curl -X POST "$API/admin/embedding/regenerate?embedding_type=all&only_incompatible=true"

The migration prints this as a NOTICE. Re-embedding is ADR-803's stated consequence of changing the universal text embedder, and it needs the API running with the new model loaded, which a schema migration cannot assume.

  1. The image loader resolves the vision tower by model_type (siglip → SiglipVisionModel, clip → CLIPVisionModel) and pools with pooler_output when the model provides one, falling back to the CLS token. SigLIP has no CLS token; its pooled output is the attention-pooling head.

  2. Task prefixes are applied as raw strings (encode(prompt=...)) rather than by registered prompt name. ModernBERT-embed registers query and document prompt names with empty strings, so prompt_name="query" would silently drop the prefix.

  3. ADR-103's "nomic-first" invariant is read as "local-first": the out-of-the-box profile is local, offline-baked, and reasoning stays remote. The specific model pair is this ADR's to choose. ADR-804 §6's model recommendation is superseded by the table above.

Consequences

Positive

  • transformers and sentence-transformers track upstream again; the last held dependency from the sweep is released.
  • No remote code in the default profile. The bake step no longer has to warm the HF_HOME/modules dynamic-module cache for offline boot.
  • Text quality is at least preserved: same dimensions, same prefixes, same context, a newer base architecture.
  • The image index moves from a 2024 vision model to SigLIP 2.

Negative

  • Every existing install on the seeded Nomic profile must re-embed concepts, sources, and vocabulary after upgrade. Search over old embeddings degrades until that runs. Cube and the development volume are both in this state.
  • The baked model layer is 2.1 GB (measured: kg-api:latest grew from 6.58 GB to 8.56 GB). SigLIP 2 ships both towers in one safetensors file, and the runtime loads only the vision tower from it.
  • The Nomic v1.5 profile row can no longer be activated: its models do not load on transformers 5.x. It is kept for history and for restores of old backups, not for use.

Neutral

  • Backups carry the embedding identity (provider:model@dims); restores of pre-081 backups report the mismatch through the existing identity check.
  • IMAGE_EMBEDDING_MODEL and the module-level defaults in embedding_model_manager.py / visual_embeddings.py follow the new pair.
  • ADR-103's Pi RAM budget line (nomic ~400MB) needs re-measuring with the new pair when Stage 2 resumes (#515).

Alternatives Considered

  • Vendor or monkey-patch the Nomic remote code. Fixing n_inner and re-adding get_extended_attention_mask in-tree is a private fork of code loaded via trust_remote_code; every transformers major would reopen it. Rejected.
  • Keep the cap. transformers 4.57 is the last 4.x line; sentence-transformers 6 requires 5. The cap would hold both indefinitely. Rejected.
  • Single multimodal SigLIP 2 profile (multimodal = true, text and image from one model). The 64-token text tower cannot embed source chunks. Rejected.
  • SigLIP 2 so400m (1152 dims, 1.1 B params) for images. Better retrieval, 2.3 GB at fp16, too heavy for the CPU appliance target. Not chosen for the default; selectable as a custom profile.
  • google/embeddinggemma-300m, Qwen/Qwen3-Embedding-0.6B for text. Both native and higher on MTEB, but 300 M and 600 M parameters against ModernBERT-embed's 149 M, and the Gemma license is not Apache 2.0. Not chosen for the default.
  • nomic-ai/nomic-embed-text-v2-moe. Remote code. Rejected for the same reason as v1.5.