AliSaadatV commited on
Commit
0a41df7
·
verified ·
1 Parent(s): 5b08fa1

Upload README.md

Browse files
Files changed (1) hide show
  1. README.md +157 -0
README.md ADDED
@@ -0,0 +1,157 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Bio-ACDC: Biological Sequence Model Coevolution
2
+
3
+ An adaptation of [AC/DC (Assessment Coevolving with Diverse Capabilities)](https://acdc-llm.github.io) for biological language models.
4
+
5
+ ## Overview
6
+
7
+ Bio-ACDC coevolves populations of biological language models (for DNA, RNA, and Protein sequences) with synthetic sequence tasks to discover specialized model experts.
8
+
9
+ ## Core Components
10
+
11
+ ### 1. Coevolution Loop (`core.py`)
12
+ - **Initialization**: Seed models are evaluated on base tasks
13
+ - **Offspring Generation**: Parents are merged and optionally mutated
14
+ - **Evaluation**: New models are tested on current task pool
15
+ - **Archive Update**: Dominated Novelty Search (DNS) maintains diverse Pareto archive
16
+ - **Task Generation**: New tasks target weaknesses discovered in the archive
17
+
18
+ ### 2. Task Pool (`tasks.py`)
19
+ Generates and manages biological sequence tasks:
20
+ - **Protein tasks**: Motif recognition, sequence completion, structure prediction
21
+ - **DNA tasks**: Regulatory element detection, motif localization
22
+ - **RNA tasks**: Secondary structure prediction, motif finding
23
+
24
+ Tasks are auto-generated targeting archive weaknesses.
25
+
26
+ ### 3. Model Merging (`mergers.py`)
27
+ Multiple merging strategies:
28
+ - **Linear Merge**: Weighted average of parent parameters
29
+ - **SLERP**: Spherical linear interpolation
30
+ - **Task Vector**: Arithmetic with base model subtraction
31
+
32
+ ### 4. Mutation (`mutators.py`)
33
+ Controlled perturbation operators:
34
+ - **Gaussian Noise**: Add random noise to weights
35
+ - **Layer Scale**: Randomly scale specific layers
36
+ - **Dropout**: Structured pruning of weights
37
+
38
+ ### 5. Archive (`archive.py`)
39
+ Dominated Novelty Search maintains a Pareto archive:
40
+ - Maximizes fitness + novelty
41
+ - Novelty based on unique capabilities vs fitter solutions
42
+ - Difficulty-aware weighting
43
+
44
+ ### 6. Evaluator (`evaluator.py`)
45
+ Evaluates models on biological tasks:
46
+ - Sequence identity/similarity
47
+ - Motif containment
48
+ - Perplexity
49
+ - RNA structure prediction
50
+
51
+ ## Usage
52
+
53
+ ### Quick Start
54
+
55
+ ```python
56
+ from bio_acdc import BioACDC, BioACDCConfig
57
+ from bio_acdc.tasks import BioTaskPool
58
+ from bio_acdc.mergers import LinearMerge
59
+ from bio_acdc.mutators import GaussianNoiseMutator
60
+ from bio_acdc.evaluator import BioEvaluator
61
+
62
+ # Configuration
63
+ config = BioACDCConfig(
64
+ seed_model_paths=[
65
+ "facebook/esm2_t33_650M_UR50D",
66
+ "InstaDeepAI/nucleotide-transformer-v2-500m-multi-species",
67
+ ],
68
+ archive_size=20,
69
+ num_generations=10,
70
+ offspring_per_gen=5,
71
+ output_dir="./bio_acdc_output",
72
+ )
73
+
74
+ # Components
75
+ task_pool = BioTaskPool(seed=42)
76
+ evaluator = BioEvaluator()
77
+ merger = LinearMerge()
78
+ mutator = GaussianNoiseMutator(std=0.01)
79
+
80
+ # Create Bio-ACDC
81
+ bio_acdc = BioACDC(
82
+ config=config,
83
+ task_pool=task_pool,
84
+ evaluator=evaluator,
85
+ merger=merger,
86
+ mutator=mutator,
87
+ )
88
+
89
+ # Run evolution
90
+ final_archive = bio_acdc.evolve()
91
+
92
+ # Best model
93
+ best = bio_acdc.archive.get_best()
94
+ print(f"Best model: {best.model_path}, Fitness: {best.fitness:.4f}")
95
+ ```
96
+
97
+ ### Using with Custom Models
98
+
99
+ ```python
100
+ # For ESM-2 protein models
101
+ config = BioACDCConfig(
102
+ seed_model_paths=[
103
+ "facebook/esm2_t33_650M_UR50D",
104
+ "facebook/esm2_t30_150M_UR50D",
105
+ ],
106
+ )
107
+
108
+ # For Nucleotide Transformer DNA models
109
+ config = BioACDCConfig(
110
+ seed_model_paths=[
111
+ "InstaDeepAI/nucleotide-transformer-v2-500m-multi-species",
112
+ "InstaDeepAI/nucleotide-transformer-v2-100m-multi-species",
113
+ ],
114
+ )
115
+ ```
116
+
117
+ ## Architecture Comparison: ACDC vs Bio-ACDC
118
+
119
+ | Feature | ACDC (SakanaAI) | Bio-ACDC |
120
+ |---------|-----------------|----------|
121
+ | Domain | General NLP (text) | Biological sequences |
122
+ | Seed Models | Qwen/Llama LLMs | ESM-2, NT, protein/DNA LMs |
123
+ | Tasks | Synthetic code/math/text | Motif detection, sequence completion, structure prediction |
124
+ | Evaluation | LLM-as-judge + code sandbox | Sequence similarity, biological metrics |
125
+ | Merging | SLERP, Linear, Task Vectors | Same (adapted for masked LMs) |
126
+ | Archive | Dominated Novelty Search | Same |
127
+ | Task Gen | LLM-based task creation | Rule-based + evolutionary targeting |
128
+
129
+ ## Key Biological Adaptations
130
+
131
+ 1. **Sequence-Specific Tasks**: Motifs, regulatory elements, structure prediction
132
+ 2. **Token-Aware Evaluation**: Handles amino acids (20 AA), nucleotides (4 NT)
133
+ 3. **Masked LM Support**: Works with ESM-2 and Nucleotide Transformer (masked language models)
134
+ 4. **Motif-Based Difficulty**: Tasks target specific biological motifs
135
+ 5. **Structure Evaluation**: RNA secondary structure comparison
136
+
137
+ ## Requirements
138
+
139
+ ```
140
+ torch>=2.0
141
+ transformers>=4.30
142
+ datasets
143
+ numpy
144
+ safetensors
145
+ biopython # Optional, for advanced sequence analysis
146
+ ```
147
+
148
+ ## Citation
149
+
150
+ ```bibtex
151
+ @software{bio_acdc_2024,
152
+ title = {Bio-ACDC: Coevolution of Biological Language Models and Sequence Tasks},
153
+ author = {Adapted from SakanaAI AC/DC},
154
+ year = {2024},
155
+ url = {https://acdc-llm.github.io}
156
+ }
157
+ ```