None
DE
Tsinghua University’s KTransformers enables full-powered DeepSeek-R1 with low-cost graphics card
['Jessie Wu', 'More Jessie Wu', '.Wp-Block-Co-Authors-Plus-Coauthors.Is-Layout-Flow', 'Class', 'Wp-Block-Co-Authors-Plus', 'Display Inline', '.Wp-Block-Co-Authors-Plus-Avatar', 'Where Img', 'Height Auto Max-Width', 'Vertical-Align Bottom .Wp-Block-Co-Authors-Plus-Coauthors.Is-Layout-Flow .Wp-Block-Co-Authors-Plus-Avatar']
TechNode
Now, with a 24GB VRAM 4090D (NVIDIA GPU), users can run the full-powered DeepSeek-R1 and V3 671B version locally.
Pre-processing speeds can reach up to 286 tokens per second, while inference generation speeds peak at 14 tokens per second.
Running the full version of DeepSeek-R1 on standard hardware is a major challenge for most users.
Details: KTransformers breaks the limitation of AI large models relying on expensive cloud servers, according to the National Business Daily report.
KTransformers optimizes the deployment of large language models (LLMs) on local machines to overcome resource limitations.
['cost'
'tokens'
'rmb'
'card'
'users'
'graphics'
'million'
'servers'
'ktransformers'
'deepseekr1'
'version'
'technode'
'models'
'enables'
'lowcost']