embedding compression
vertebrae can apply an optional embedding-compression step after feature
extraction and before OverlapIndex scoring, stability analysis, Separatix
diagnostics.
This is useful when you want to:
compare raw embeddings against compressed variants,
test whether a lower-dimensional representation preserves class separation,
evaluate sparse text embeddings with
TruncatedSVD,study matryoshka-style prefix truncation,
measure the effect of lossy precision reduction such as
float16orint8.
Compression runs after extraction rather than inside the extractor. When the raw
embedding is cache-eligible, it remains independently reusable and each compression
recipe receives its own derived artifact key and report metadata. Derived eligibility
is never broader than source eligibility: CacheConfig(enabled=False), hosted
cache opt-out, or cache_status="bypassed_unsafe_identity" also prevents compressed
results from being reused. The compression may still run for the current evaluation;
it simply is not published as a reusable cache artifact.
Configuration
Use EmbeddingCompressionConfig with Evaluator or Benchmark:
from vertebrae import BenchmarkDataset, EmbeddingCompressionConfig, Evaluator, DatasetIdentity
from vertebrae.config import CacheConfig, StabilityConfig
from vertebrae.extractors import PrecomputedExtractor
dataset = BenchmarkDataset.from_embeddings(embeddings=Z, labels=y, identity=DatasetIdentity.declared("example-dataset", "1"))
compression = EmbeddingCompressionConfig(
enabled=True,
method="pca",
n_components=128,
)
result = Evaluator(
dataset=dataset,
extractor=PrecomputedExtractor(name="baseline"),
compression_config=compression,
stability_config=StabilityConfig(enabled=False),
cache_config=CacheConfig(cache_dir=".vertebrae_cache"),
).run()
To compare multiple compression variants in one run, pass
compression_configs=[...] to Benchmark.
For dimensional compression methods, dtype may be set to a real floating-point
NumPy dtype such as float16, float32, or float64. Integer, complex, boolean,
object, text, datetime, and structured dtypes are rejected. Use quantization’s
precision="int8" or precision="uint8" when evaluating integer encoding.
Supported methods
none
Disables compression and preserves the current workflow.
pca
Dense PCA for low- to medium-dimensional dense embeddings.
Requires dense input.
Requires exactly one of
n_componentsorpreserve_variance.preserve_variancemust be strictly between0and1.Supports
whiten=True.
incremental_pca
Dense PCA variant for larger dense embeddings.
Requires dense input.
Requires
n_components.Supports
whiten=True.
truncated_svd
Recommended for sparse text embeddings such as TF-IDF features.
Accepts sparse input.
Requires
n_components.
gaussian_random_projection
Fast dense or sparse projection for coarse dimensionality reduction.
Accepts sparse input.
Requires
n_components.
sparse_random_projection
Sparse-friendly random projection for very high-dimensional sparse matrices.
Accepts sparse input.
Requires
n_components.
prefix_truncate
Keeps the first n_components dimensions without fitting a model.
Accepts dense and sparse input.
Requires
n_components.Honors
dtypeafter truncation while preserving sparse storage.Sparse
float16output is rejected because SciPy does not support it consistently across versions; use sparsefloat32orquantizewithprecision="float16"instead.Use
assume_matryoshka=Truewhen the embedding source is intentionally dimension-ordered, such as matryoshka-trained or shortened embeddings.
When assume_matryoshka=False, vertebrae records a warning so reports make it
clear this is being treated as a prefix-based diagnostic rather than a claimed
model-native shortening path.
For every dimensional method, a requested n_components greater than or equal to
the current embedding width is recorded as a skipped compression and leaves values,
sparsity, dimensions, and dtype unchanged. A requested dtype is applied only when
compression is actually performed. It is never used to cast a skipped or disabled
compression result. PCA and incremental PCA also reject a reducing dimension that
exceeds the number of samples on the side used to fit the transform.
quantize
Applies a lossy precision-reduction step and returns numeric embeddings that can still be scored by OverlapIndex.
Supported precisions:
float16: direct cast for dense embeddings; sparse values are rounded throughfloat16and returned as a sparsefloat32scoring matrix.int8: dense scalar quantize/dequantize round trip using symmetric per-dimension scaling.uint8: dense scalar quantize/dequantize round trip using affine per-dimension min/max scaling.
The integer precisions and sparse float16 record the encoded dtype and estimated
encoded size in metadata, then return float32 embeddings for scoring. They do
not return unsupported or integer matrices to the scoring pipeline.
Binary and packed-bit quantization are intentionally not included here because they are more appropriate for retrieval or ANN index evaluation than for MiniBatchKMeans-backed OverlapIndex scoring.
Metadata and reports
Each extractor result includes compression_metadata describing the applied
compression. Depending on the method, this may include:
methodprecisionoriginal_dimcompressed_dimdtypecompression_ratioexplained_variance_totalassume_matryoshkaquantization calibration metadata
warnings
Markdown reports and result.to_dataframe() include compression columns so raw
and compressed variants can be compared side by side.
CLI
The artifact-based CLI also supports compression:
vertebrae compress \
--cache-dir .vertebrae_cache \
--embedding-key embeddings/... \
--method prefix_truncate \
--n-components 256 \
--assume-matryoshka
For PCA, pass either --n-components or --preserve-variance; the CLI rejects
commands that provide both.
Example quantization run:
vertebrae compress \
--cache-dir .vertebrae_cache \
--embedding-key embeddings/... \
--method quantize \
--precision int8
The resulting compressed artifact can be passed to vertebrae score just like a
raw embedding artifact.
Choosing a method
Use
truncated_svdfor sparse text features.Use
pcafor dense embeddings when you want an interpretable learned reduction.Use
prefix_truncatefor matryoshka-style or dimension-shortened embeddings.Use
quantizewhen you want to see how precision reduction affects the diagnostic, not when you need a retrieval index benchmark.