AI on Trial: Fine-Tuning LLMs to Judge other LLMs; data, results and challenges
Key take-aways (internal notes)
- directly training on preference data, even if low signal, improves the performance as an evaluator
- it’s inconclusive if first generating the feedback and then the score improves the performance. Hypothesis: the criteria for the preference dataset is vague, thus, a CoT might not be helpful
- dataset biases such as length, position or score pose challenges for training a robust and unbiased evaluator
Motivation (1’)
Experiments (5’)
Challenges (2’)
Future Experiments (1’)