LoRA and parameter efficient fine tuning
Fine tune large models on modest hardware by training small low rank adapters instead of every weight
A taste of a lesson
I set LoRA rank 64 instead of 8 and my results got no better. Did I do something wrong?
Probably not. Rank sets the capacity of the update, and many tasks only need a small change in behaviour, so extra rank adds parameters without adding useful signal. If quality is flat, the bottleneck is usually elsewhere: data quality, which layers you target, or the learning rate. Two quick checks: confirm the trainable parameter count actually rose about eightfold, and try adding the feed forward layers as targets at rank 8. Also remember alpha over r scales the update, so changing r alone can change the effective step size. What learning rate and alpha did you use for each run?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain the low rank update B A and how rank and alpha shape it
- Count trainable parameters and estimate memory savings for an adapter setup
- Choose target layers, rank and learning rate as testable starting points
- Decide whether to merge adapters or serve them separately
- Judge when full fine tuning is likely to be the better option
Lesson plan
- 1 Why full fine tuning is expensive See where memory goes in full fine tuning and what freezing weights saves. Start
- 2 The low rank update Understand how B times A approximates a full weight change with far fewer numbers. Start
- 3 Choosing rank, alpha and target layers Pick sensible starting settings and know how to test alternatives. Start
- 4 Training on a quantised base Fit bigger models by storing frozen weights in low precision while adapters train. Start
- 5 Merging, swapping and serving adapters Deploy adapters either merged into the base or as interchangeable modules. Start
- 6 Other methods and when to go full Compare LoRA with other parameter efficient methods and with full fine tuning. Start
Try asking
About this tutor
For learners who understand basic fine tuning and want to adapt large models without the memory bill of full fine tuning. You will work out how a low rank update replaces a full weight change, count the parameters it saves, and learn what rank, alpha and target layers actually control. Then you cover the practical side: training adapters on a quantised base model, choosing a learning rate, merging adapters into the base weights or keeping them separate to serve many tasks from one model. The course ends with other parameter efficient methods and an honest look at when full fine tuning is still the better choice.
Reviews
4.7
3 ratingsSample
- Tobias R.Sample
Printing the trainable parameter count caught my bug on day one: my target layer names did not match the model, so only a handful of weights were training.
- Lucas F.Sample
We now serve three customer adapters on one base model. The merging versus swapping lesson helped me explain the trade off to my team.
- Anjali M.Sample
The parameter counting lesson made the whole method click. I wish there had been a bit more on choosing alpha, but the advice to test two settings was sensible.
About the teacher
Fine tuning with judgment: when to do it, how to do it well, and how to know it worked
9 tutors 428 lessons taught Sample
I teach fine tuning and post training: choosing between prompting, retrieval and tuning, building datasets, parameter efficient methods, instruction and preference tuning, and evaluating the result. My background is in applied machine learning projects where the expensive mistake was usually tuning a model before anyone had defined what better meant. That is why I start every topic with the evaluation...
See Neha's profile and tutorsMore like this
Other tutors on the same or nearby topics.