Investigating Model Compression Techniques for Federated Learning in LLMs within Trained Environments
2025 (English)Independent thesis Advanced level (degree of Master (One Year)), 10 credits / 15 HE credits
Student thesis
Abstract [en]
Large Language Models (LLMs) like T5-small offer powerful natural language understanding and generation capabilities but present significant challenges for deployment in resource-constrained and privacy-sensitive Federated Learning (FL) environments. This thesis investigates the effectiveness of combining parameter-efficient fine-tuning (PEFT) techniques, specifically Low-Rank Adaptation (LoRA), with other model compression strategies such as pruning and knowledge distillation, to enable efficient FLof the T5-small model. The study evaluates how these combined approaches affect the balance between model performance (measured by loss and perplexity) and resource utilisation (processing time, peak memory consumption) within a simulated FL environment employing the FedAvg algorithm on the HealthCareMagic dataset.Three main configurations were compared: a standard Supervised Fine-tuned (SFT) baseline, a LoRA-PEFT model (utilising LoRA), and an aggressively compressed model incorporating LoRA, base model pruning, and distillation. Experimental results demonstrate that LoRA-PEFT significantly reduced total processing time (e.g., by approximately 13.5% compared to the SFT baseline) while maintaining a comparable peak memory footprint during FL training when operating on a quantized base. However, this efficiency was accompanied by a notable degradation in predictive performance (higher loss and perplexity). The introduction of pruning and distillation to the LoRA model further reduced processing time (achieving the fastest overall time) but resulted in the most significant compromise in model accuracy. The SFT baseline, while the most computationally intensive, achieved the best performance. This research highlights the critical trade-offs inherent in applying compression techniques to LLMs within federated settings. While PEFT and aggressive compression can substantially improve computational efficiency, they can also impact model fidelity. The findings offer insights into the practical considerations for deploying LLMs in resource-limited scenarios, emphasizing the need for careful calibration of compression strategies to balance performance requirements with efficiency gains. Trained model artifacts and a demonstration application have been made publicly available via the Hugging Face Hub and Spaces.
Place, publisher, year, edition, pages
2025. , p. 58
Keywords [en]
Federated Learning, Large Language Models, T5-small, Parameter- Efficient Fine-Tuning, LoRA, Model Compression, Pruning, Knowledge Distillation, Resource-Constrained AI, Healthcare AI
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:his:diva-25693OAI: oai:DiVA.org:his-25693DiVA, id: diva2:1986644
Subject / course
Informationsteknologi
Educational program
Data Science - Master’s Programme
Supervisors
Examiners
2025-08-012025-08-012025-09-29Bibliographically approved