Dataset and Qwen2.5-3B adapters (SFT → GRPO) for multi-turn mystery investigation with tools, unlock DAGs, and cited case closes.
Vaidik
VaidikML0508
AI & ML interests
exploring another way to use gradient decent
Recent Activity
liked a model 1 day ago
VaidikML0508/qwen2.5-3b-investigation-grpo upvoted a collection 6 days ago
Mystery Investigation Agent updated a model 6 days ago
VaidikML0508/qwen2.5-3b-investigation-sftOrganizations
None yet
Shark Tank Deal Evaluator
This collection features Llama-3.2-3B models fine-tuned to simulate Shark Tank deal evaluations and decision-making based on company pitches and offer
-
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-GRPO-16bits-V1
Text Generation • 3B • Updated • 7 • 1 -
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-SFT-DPO-4bits-V1
Text Generation • 3B • Updated • 8 -
VaidikML0508/SharkTank-Offer-V1
Viewer • Updated • 255 • 13 -
VaidikML0508/SharkTank-Offer-DPO-dataset-V1
Viewer • Updated • 263 • 16 • 1
Mystery Investigation Agent
Dataset and Qwen2.5-3B adapters (SFT → GRPO) for multi-turn mystery investigation with tools, unlock DAGs, and cited case closes.
Shark Tank Deal Evaluator
This collection features Llama-3.2-3B models fine-tuned to simulate Shark Tank deal evaluations and decision-making based on company pitches and offer
-
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-GRPO-16bits-V1
Text Generation • 3B • Updated • 7 • 1 -
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-SFT-DPO-4bits-V1
Text Generation • 3B • Updated • 8 -
VaidikML0508/SharkTank-Offer-V1
Viewer • Updated • 255 • 13 -
VaidikML0508/SharkTank-Offer-DPO-dataset-V1
Viewer • Updated • 263 • 16 • 1
models 16
VaidikML0508/qwen2.5-3b-investigation-sft
Text Generation • Updated • 23
VaidikML0508/qwen2.5-3b-investigation-grpo
Text Generation • Updated • 27 • 1
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-GRPO-16bits-V1
Text Generation • 3B • Updated • 7 • 1
VaidikML0508/Shark-Tank-Offer-Evaluator-llama3.2-3B-Instruct-SFT-DPO-4bits-V1
Text Generation • 3B • Updated • 8
VaidikML0508/rl_course_vizdoom_health_gathering_supreme
Reinforcement Learning • Updated
VaidikML0508/Reinforce-pixel-copte-1
Reinforcement Learning • Updated
VaidikML0508/Reinforce-pixel-copter
Updated
VaidikML0508/ML-Agents-Pyramids
Reinforcement Learning • Updated • 1
VaidikML0508/ppo-LunarLander-v2
Reinforcement Learning • Updated • 3
VaidikML0508/taxi-V3
Reinforcement Learning • Updated