That’s what Duality, a company specializing in privacy-enhancing technologies, hopes to accomplish with its private large language model (LLM) inference framework. “Fully homomorphic encryption algorithms are heavily memory bound,” says Rashmi Agrawal, cofounder and chief technology officer at CipherSonic Labs, a company that spun out of her doctoral research at Boston University on accelerating homomorphic encryption. Hardware acceleration also plays a part in speeding up FHE for LLM inference, especially for bigger AI models. She adds that FPGAs in particular can be tailored for fully homomorphic encryption workloads. For Duality’s next steps, the team is progressing their private LLM inference framework from prototype to production.