On Tuesday, Chinese GPU (Graphics Processing Unit) maker Moore Threads announced the rapid deployment of DeepSeek’s distilled model inference services, enabling large-scale model capabilities to be transferred to smaller, more efficient versions for high-performance inference on domestic GPUs. Details: Based on its self-developed GPU, Moore Threads has quickly deployed inference services for the DeepSeek distilled model through a dual-engine approach that combines open-source and proprietary technologies. According to the announcement, Moore Threads has deployed the DeepSeek-R1-Distill-Qwen-7B distilled model based on the Ollama open-source framework. This engine ensures the efficient operation of DeepSeek’s distilled model while laying a strong foundation for future large-scale model deployments, the company added. Moore Threads said it aims to empower more developers in AI innovation using its GPU by deploying inference services for the DeepSeek distilled model.