Spaces:
Running
Running
Update README.md
Browse files
README.md
CHANGED
|
@@ -20,7 +20,7 @@ equal participation, 2) administer the test among council members, and 3) evalua
|
|
| 20 |
responses as a collective jury.
|
| 21 |
|
| 22 |
Our initial research deploys a council of 20 newest LLMs on an open-ended emotional intelligence task: responding to interpersonal dilemmas. Our results show that the LMC produces rankings that are more separable, robust,
|
| 23 |
-
and less biased than those from any individual LLM judge, and is more consistent with a human-established leaderboard compared to other benchmarks.
|
| 24 |
|
| 25 |
Roadmap:
|
| 26 |
|
|
|
|
| 20 |
responses as a collective jury.
|
| 21 |
|
| 22 |
Our initial research deploys a council of 20 newest LLMs on an open-ended emotional intelligence task: responding to interpersonal dilemmas. Our results show that the LMC produces rankings that are more separable, robust,
|
| 23 |
+
and less biased than those from any individual LLM judge, and is more consistent with a human-established leaderboard compared to other benchmarks like Chatbot Arena or MMLU.
|
| 24 |
|
| 25 |
Roadmap:
|
| 26 |
|