Högskolan i Skövde

his.sePublications
4243444546474845 of 356
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • apa-cv
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Evaluating the impact of chunking strategies and embedding models on retrieval performance in naive RAG
University of Skövde, School of Informatics.
University of Skövde, School of Informatics.
2026 (English)Independent thesis Basic level (degree of Bachelor), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Large Language Models (LLMs) have advanced Natural Language Processing (NLP) but remain prone to hallucinations, particularly in domain-specific applications such as legal text processing. Retrieval-Augmented Generation (RAG) addresses this limitation by grounding responses in retrieved documents. This study investigates the main and interaction effects of chunking strategies and embedding models on retrieval performance in a Naive RAG pipeline using a GDPR-based legal corpus. A controlled factorial experiment evaluated recursive, token-based and sentence-based chunking in combination with the OpenAI text-embedding- 3-small and text-embedding-3-large models using Recall@k, Mean Reciprocal Rank (MRR) and normalized Discounted Cumulative Gain (nDCG@k). The results show that both the chunking strategy and the embedding model have statistically significant effects on retrieval performance, with the chunking strategy being the dominant factor. A statistically significant interaction effect was also detected, although the effect sizes for both the main and interaction effects were small. Sentence-based chunking achieved the strongest overall performance, while token-based chunking produced the highest MRR. Inferential analyses indicate that these differences should be interpreted with caution, given their limited practical significance. 

Place, publisher, year, edition, pages
2026. , p. ii, 45, v
Keywords [en]
Large Language Models (LLMs), Natural Language Processing (NLP), Retrieval-Augmented Generation (RAG), Information Retrieval (IR), Chunking Strategy, Embedding model, General Data Protection Regulation (GDPR)
National Category
Information Systems, Social aspects
Identifiers
URN: urn:nbn:se:his:diva-26855OAI: oai:DiVA.org:his-26855DiVA, id: diva2:2084161
Subject / course
Informationsteknologi
Educational program
Computer Science - Specialization in Systems Development
Supervisors
Examiners
Available from: 2026-07-03 Created: 2026-07-03 Last updated: 2026-07-03Bibliographically approved

Open Access in DiVA

fulltext(1780 kB)20 downloads
File information
File name FULLTEXT01.pdfFile size 1780 kBChecksum SHA-512
c7422b0902f20cad54ca9634ab112334402adafc60076833767bff66b86c7de93266d673c57befc8c75942441ef484e9df5019a37f3b35b9f2fdd7465c36e009
Type fulltextMimetype application/pdf

By organisation
School of Informatics
Information Systems, Social aspects

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 86 hits
4243444546474845 of 356
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • apa-cv
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf