Sep 20264 min read

Fine tuning

Fine-TuningLLMs

A couple years ago, I had tried to fine tune Llama 8B with a dataset of my conversations with people on colab. I had given up on the idea because training anything on colab was a pain. You had to keep your computer on for the entirety of the training process and the GPU disconnected sometimes.

A few weeks ago I found out about modal and its serverless model and most importantly that it gave $30 every month to its users free of cost. I had to try again. I had some 100k question-answer pairs that I had gathered from Whatsapp, Instagram and Signal. I had learnt from earlier that most text messages are not profound and that I was an unusually dry texter. So i trimmed down the dataset to 30k long-ish messages with a few short messages here and there. I trained it for 4 epochs on an Nvidia L4. That took about 10 hours.

After testing, I found out the model had severely overfitted. Also, I had trained the model incorrectly. I had computed the loss on the other person's response too, so the model was learning on their parts of the conversation as well; Then, I forgot to add the <|header_id> tag that lets the model know whose turn it is to speak. So, it started speaking for both people. On top of that, I realized that 95% of my data came from conversations with one person. So, sometimes it would just blurt out exact phrases from my conversation with this person. This was bad.

Maybe I was training on too much data and for too long. I tried a few other approaches. I changed models, tried tweaking hyperparameters, trained it on some long form content I'd written. Lastly, I tried training Qwen 14B on 1.5k pairs of just funny things I'd said. I did this by filtering texts that were followed by laughing emojis and other things. Did this mean i was funny only 1.5% of the time? Remember that the data depends on the other person laughing or even understanding all of my other comedic masterpieces. The actual number is far higher I'd like to believe. During the second world war, allies decided to add more armor to the parts of the plane that were the most damaged of the planes that returned. They failed to account for the planes that had never returned. I believe a similar thing is true for my funny percentage.

Out of all of my attempts, Qwen trained on that 1.5k dataset was the most fun. It said the most absurd yet hilarious things. I was planning on giving a few friends access to these models. In fact, I had deployed a whole backend in Go on a tiny server that would let authenticated users talk to my models by waking up a GPU on Modal everytime someone makes a request. Talking to any of the models for just five minutes made me realize that I cannot trust other people with these models. Since they were trained on my own conversations, sometimes they would say things that I would not want people to know and other times they would just hallucinate. When I asked about my celebrity crush, Qwen answered Karan Johar.

qwen-14b[sidebar]
You: go is faster
Aryan (qwen-comedy): Doesn't matter really
You: only thing that matters
Aryan (qwen-comedy): For backend
You: speed and concurrecy is the only thing that matters
Aryan (qwen-comedy): If you're doing backend
You: so what would you use python for
Aryan (qwen-comedy): Web scraping
Data analysis
You: go is objectively better for web scraping
Aryan (qwen-comedy): It's fine for backend stuff
You: you have no argument
Aryan (qwen-comedy): Yes I know
You: your're just making me angry
Aryan (qwen-comedy): I know
You: then why are you talking to me
Aryan (qwen-comedy): You were talking to me first

Training LLMs is fun. 10/10. Highly recommend.