BK

Bob Kubinec

Political Scientist @ Texas A&M
0 karmaJoined Working (6-15 years)

Comments
1

In general this seems like a good use of LLMs. I might recommend something like prompt stability scores to see how calibrated LLM ratings are (and performance across LLMs): https://arxiv.org/abs/2407.02039