Time-Embed BGE-M3 LMS Temporal v2.0 Multisource

한국어 LMS 검색에서 상대 시간 표현과 검색 의도를 함께 구분하도록 BAAI/bge-m3를 미세조정한 연구 모델입니다. 모델 구조와 손실 함수는 변경하지 않고 학습 데이터 구성을 개선했습니다.

Training data

총 234,296개 grouped training row를 사용했습니다.

  • Korean temporal inventory C5: 120,000
  • Real-query temporal LMS data: 54,296
  • LMS intent-control replay: 60,000

각 학습 step은 query 1개, positive 1개, negative 7개(train_group_size=8)를 사용합니다. 평가 Dev/Test의 문장 및 row ID와 학습 데이터의 중복은 모두 0건입니다.

Training

  • Base model: BAAI/bge-m3
  • Selected checkpoint: step 1250
  • Learning rate: 1e-6
  • Per-device batch size: 1
  • Gradient accumulation: 32
  • Precision: BF16
  • Gradient checkpointing: enabled
  • Seed: 42
  • Hardware: RTX 4070 Ti SUPER 16GB

Results

Temporal Dev

Metric Result
Pairwise accuracy 0.951020
Margin p10 0.109466
Hard-negative violation rate 0.048980
Random-pair p95 0.576171
Near-one rate 0.014184

One-time frozen Inventory Test

Metric Result
Pairwise accuracy 0.972028
Margin mean 0.293959
Margin p10 0.115584
Hard-negative violation rate 0.027972
Random-pair p95 0.542047
Near-one rate 0.013530

Semantic retention versus Base

  • LMS intent delta: +0.159000
  • KLUE-STS delta: +0.003883
  • KorSTS delta: +0.011948
  • KLUE-NLI delta: +0.030642
  • SQuADKorV1 Retrieval delta: -0.013460

Usage

from FlagEmbedding import BGEM3FlagModel

model = BGEM3FlagModel(
    "kev-KOH/time-embed-bge-m3-lms-temporal-v2-0-multisource",
    use_fp16=True,
)
embeddings = model.encode(
    ["일주일 전에 올라온 강의자료 찾아줘", "7일 전에 등록된 자료 보여줘"]
)["dense_vecs"]

This repository contains model weights and tokenizer/configuration files only. Frozen test examples and private evaluation artifacts are not included.

Downloads last month
20
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kev-KOH/time-embed-bge-m3-lms-temporal-v2-0-multisource

Base model

BAAI/bge-m3
Finetuned
(571)
this model