Researchers at KAIST have created PIMBA, a new semiconductor chip that acts as a brain for advanced artificial intelligence (AI) models. This chip could solve big problems in AI by making it run four times faster while using 2.2 times less power. AI models now handle long and complex sentences. But as these models grow bigger, they need more speed and less energy.
The chip uses a hybrid structure. It combines Transformer and Mamba. Transformer is the main method in today's AI tools like ChatGPT. It looks at all words in a sentence at once, which is powerful but slow and energy-hungry for long text. Mamba is a newer design that processes words one by one over time, saving energy but still facing limits. PIMBA blends the best of both to get speed and efficiency.
Processing-in-Memory
PIMBA works with Processing-in-Memory, or PIM. This means the chip does math right inside its own storage, where data lives. In normal chips like GPUs data must move out of storage to a separate area for calculations. Moving data wastes time and power. PIMBA skips this move. It calculates everything in place.
Tests show PIMBA boosts processing speed by up to 4.1 times. It cuts energy use by 2.2 times on average compared to GPUs. KAIST's School of Computing led the work with partners from Georgia Institute of Technology in the US and Uppsala University in Sweden.
The researchers will present the methods used in PIMBA's developments and preliminary results on October 20 at MICRO 2025, a conference on computer design in Seoul. The paper already won a gold prize from Samsung. Support came from South Korea's government programs for AI chips and ICT research, plus tools from the IC Design Education Center. This chip could make AI devices smaller, cheaper, and kinder to the environment for phones, servers, and robots.