justinxzhao commited on
Commit
441a25f
·
verified ·
1 Parent(s): f6a9450

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +19 -1
README.md CHANGED
@@ -7,4 +7,22 @@ sdk: static
7
  pinned: false
8
  ---
9
 
10
- Edit this `README.md` markdown file to author your organization card.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7
  pinned: false
8
  ---
9
 
10
+ <p align="center">
11
+ <img src="https://cdn-uploads.huggingface.co/production/uploads/6462ac71514ee1645bd1f7f7/6MkoY412i9IqvISWSS4qs.png">
12
+ </p>
13
+
14
+ The rapid advancement of Large Language Models (LLMs) necessitates robust
15
+ and challenging benchmarks.
16
+
17
+ To address the challenge of ranking LLMs on *highly subjective* tasks such as emotional intelligence, creative writing, or persuasiveness,
18
+ the **Language Model Council (LMC)** operates through a democratic process to: 1) formulate a test set through
19
+ equal participation, 2) administer the test among council members, and 3) evaluate
20
+ responses as a collective jury.
21
+
22
+ Our initial research deploys a council of 20 newest LLMs on an open-ended emotional intelligence task: responding to interpersonal dilemmas. Our results show that the LMC produces rankings that are more separable, robust,
23
+ and less biased than those from any individual LLM judge, and is more consistent with a human-established leaderboard compared to other benchmarks.
24
+
25
+ Roadmap:
26
+
27
+ - Expand to more domains, use cases, and sophisticated agentic interactions.
28
+ - Produce a generalized user interface for Council-as-a-Service.