B · Normal
[Yuannao Server launches G3.5 layer all-flash solution that supports Mooncake KV cache sharing] On September 14th, Yuannao Server launched a G3.5 layer all-flash solution that supports Mooncake KV cache sharing for large model inference scenarios. The solution uses the Yuannao all-flash server NF5286 as the core, and cooperates with the Mooncake cache component and the cache-aware scheduling of the inference framework to provide the inference cluster with a KV Cache shared space that is independent of the computing nodes and can be planned individually according to business needs. When the context demand increases, the cache resources can be expanded independently; when the GPU node is expanded, adjusted, or maintained, the cache resources do not have to completely change with the computing node.
Comments