FlashAccel利用HBF提升LLM推理吞吐量
原文:{INTRESTING PAPER BASED ON HBF}2607.10186] FlashAccel: Leveraging High-Bandwidth Flash (HBF) for High-Throughput LLM Inference
HBF gives 8x - 16x more capacity than HBM at same cost, and with bandwidth till 3 tb/s. submitted by /u/9r4n4y [link] [comments]