Tag index

#Chunking

03 entries
№003 embedding-research-engineering · 09

긴 문서 Embedding: Chunking·Contextual·Multi-vector (9/14)

Chunk-then-embed의 문맥 손실부터 long-context single vector, late chunking, CDE·situated embedding, token multi-vector까지 비교하고 긴 문서 retrieval 평가를 설계합니다.

#Embeddings #LongContext #Chunking #MultiVector
긴 문서를 독립 chunk, 전체 문맥을 본 contextual chunk, single document vector와 token multi-vector로 표현하는 네 가지 retrieval 경로
№001 rag-retrieval-foundations · 02

RAG Chunking: 크기·Overlap·Semantic·Late Chunking 선택법 (2/10)

고정 길이·overlap·문서 구조·semantic·late chunking의 원리와 비용을 비교하고, tokenizer 기반 구현·parent-child 연결·평가 grid를 통해 자신의 문서와 질문에 맞는 chunk 경계와 크기를 선택하는 방법을 배웁니다.

#RAG #Chunking #Retrieval #Embedding
하나의 문서를 fixed token, structure-aware, semantic, late chunking으로 나누고 검색용 child와 답변용 parent를 연결하는 비교