_mean_pool(token_embeddings, attention_mask) reads list(_model_pooling_strategies.values())[-1] — the last-inserted value — instead of the strategy belonging to the model whose embeddings are being pooled. This works today because production loads one encoder per process and every test populates at most one dict entry, but breaks silently the moment two backbones share a process. Fix: thread model_id through _mean_pool and its call sites (_build_description_embeddings and _embed_task), resolving the strategy via _model_pooling_strategies.get(model_id, 'mean') instead of reading the global tail. Add test_mean_pool_dispatches_by_model_id_not_load_order which proves the failure mode is closed: two models with different strategies (cls, mean) are each pooled correctly regardless of which was inserted first.
52 KiB
52 KiB