№003 embedding-research-engineering · 09
긴 문서 Embedding: Chunking·Contextual·Multi-vector (9/14)
Chunk-then-embed의 문맥 손실부터 long-context single vector, late chunking, CDE·situated embedding, token multi-vector까지 비교하고 긴 문서 retrieval 평가를 설계합니다.
Tag index
Chunk-then-embed의 문맥 손실부터 long-context single vector, late chunking, CDE·situated embedding, token multi-vector까지 비교하고 긴 문서 retrieval 평가를 설계합니다.
문서 수집과 chunking부터 embedding·cosine retrieval·context 조립·근거 기반 생성·Recall@k 평가까지 최소 RAG를 Python으로 연결하고, 실패를 검색과 생성 단계로 분리해 추적합니다.
고정 길이·overlap·문서 구조·semantic·late chunking의 원리와 비용을 비교하고, tokenizer 기반 구현·parent-child 연결·평가 grid를 통해 자신의 문서와 질문에 맞는 chunk 경계와 크기를 선택하는 방법을 배웁니다.