№004 embedding-research-engineering · 12
Embedding 압축과 ANN 서빙: Dimension부터 Re-index까지 (12/14)
Matryoshka 차원 축소, float16·int8·binary·PQ 양자화, HNSW·IVF·DiskANN 검색을 분리 평가하고 품질·메모리·지연·재색인의 Pareto frontier를 설계합니다.
Tag index
Matryoshka 차원 축소, float16·int8·binary·PQ 양자화, HNSW·IVF·DiskANN 검색을 분리 평가하고 품질·메모리·지연·재색인의 Pareto frontier를 설계합니다.
Candidate 수와 token 길이로 용량을 계산하고 dynamic batching, TEI·vLLM, quantization parity를 설계합니다. Adaptive cascade, timeout·fallback·관측성으로 p95 SLO를 지킵니다.
LLM 추론의 FP8·INT8·INT4·FP4 표현, scale 단위, calibration과 outlier 처리를 비교합니다. Weight·activation·KV를 줄인 뒤 fused kernel, 실제 byte, 정확도와 SLO goodput으로 빠른 경로를 검증합니다.
LLM 양자화를 weight-only, weight·activation, KV cache로 나누고 scale·zero point·granularity를 계산하며 GPTQ, AWQ, SmoothQuant, FP8, KIVI의 memory·kernel·품질을 설계합니다.