In general this seems like a good use of LLMs. I might recommend something like prompt stability scores to see how calibrated LLM ratings are (and performance across LLMs): https://arxiv.org/abs/2407.02039
In general this seems like a good use of LLMs. I might recommend something like prompt stability scores to see how calibrated LLM ratings are (and performance across LLMs): https://arxiv.org/abs/2407.02039