justinxzhao commited on
Commit
b286409
·
verified ·
1 Parent(s): 441a25f

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -20,7 +20,7 @@ equal participation, 2) administer the test among council members, and 3) evalua
20
  responses as a collective jury.
21
 
22
  Our initial research deploys a council of 20 newest LLMs on an open-ended emotional intelligence task: responding to interpersonal dilemmas. Our results show that the LMC produces rankings that are more separable, robust,
23
- and less biased than those from any individual LLM judge, and is more consistent with a human-established leaderboard compared to other benchmarks.
24
 
25
  Roadmap:
26
 
 
20
  responses as a collective jury.
21
 
22
  Our initial research deploys a council of 20 newest LLMs on an open-ended emotional intelligence task: responding to interpersonal dilemmas. Our results show that the LMC produces rankings that are more separable, robust,
23
+ and less biased than those from any individual LLM judge, and is more consistent with a human-established leaderboard compared to other benchmarks like Chatbot Arena or MMLU.
24
 
25
  Roadmap:
26