Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Andy Chen's picture
๐Ÿ”„ In a Training Loop

Andy Chen

andynoodles
14 34
dipankarsarkar's profile picture
ยท
  • yi-hsiang-chen-tw

AI & ML interests

Feel free to contact me through Linkedin

Recent Activity

repliedto onekq's post 4 days ago
My take on device-side inference: it's all about high bandwidth memory (thinking about it, this holds for the cloud too). MacBooks enjoy incidental capacity of apple silicon, but per-device RAM is too low (16 to 24GB), only sufficient for a decent SLM. 512GB is the highest you can go (Kimi K2*). Counting MLX downloads of Kimi K2* on Huggingface, I estimate the user base to be <25K. On the other hand, the newly debuted DGX station (Nvidia) has 748GB, which can fit in the latest Kimi, DS, and Qwen. Also the quantization options of CUDA is way better than MLX. For high-end inferencing, I place my bet on workstations over Macs.
liked a model 4 days ago
zai-org/GLM-5.3
liked a model 4 days ago
zai-org/GLM-5.3-Flash
View all activity

Organizations

andynoodlesORG's profile picture andynoodlesORG2's profile picture

andynoodles 's Spaces 1

Sleeping
Agents

CloudOrAPI

๐Ÿ‘

Compare Cloud compute bills vs API bills

May 4
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs