-
VectorFLASH: A 2.8-8.7 Tokens/s Resource-Efficient Long-Context LLM Accelerator with Dual Vector Qua…
Jungwan Lee, Seongyon Hong, Youngjin Moon, Jingu Lee, Junhyuk Lee, Yuseon Choi, Jungjun Oh, and Hoi-Jun Yoo
-
NeuLPU: A 3.6 mW Spiking LLM Processing Unit for On-Device Long-Context Inference
Sangmyoung Lee, Sangwoo Ha, Youngjin Moon, Minsung Kim, Seryeong Kim, Wooyoung Jo, Junhyuk Lee, Sangyeob Kim, and Hoi-Jun Yoo
-
Trusted-Agent: A 0.535–0.539 mJ/Token Energy-Efficient Trusted Neural Processor for On-Device Person…
Wonhoon Park, Sanghyuk An, Hongseok Lee, Minsung Kim, Seryeong Kim, Jiwon Choi, Junha Ryu, Junhyuk Lee, Minhan Kang, Dongseok Im, and Hoi-Jun Yoo