New GPU architecture set to have ten times the performance of Maxwell – for deep learning.

If you thought NVIDIA’s Maxwell architecture is awesome, then GPUs developed on the Pascal architecture will surely blow your mind. Expected to be ready next year, Pascal-based GPUs is set to accelerate “deep learning applications” by 10 times the current speed Maxwell processors are capable of.
NVIDIA CEO, Jen-Hsun Huang, announced Pascal and the new NVIDIA processor’s roadmap at a keynote during the GPU Technology Conference. He said: “It will benefit from a billion dollars’ worth of refinement because of R&D done over the last three years,” in reference to Pascal.
Due to deep learning, a process where computers use neural networks to learn from one another, NVIDIA decided to modify the design of Pascal. Now it features three main features that ensures much faster and accurate learning for such networks.
Over two times the memory capacity of the GTX Titan X at 32GB, Pascal will have mixed-precision computing, 3D memory, and NVLINK. These features combine to deliver the 10 times speed improvement mentioned above. Below are brief summaries of the features as written on NVIDIA’s blog:
Mixed-Precision Computing – for Greater Accuracy
– Computes at 16-bit floating point accuracy twice as fast when compared to 32-bit floating point accuracy, which helps the classification and convolution part of deep learning.
3D Memory – for Faster Communication Speed and Power Efficiency
– Three times the bandwidth and almost three time the frame buffer capacity of Maxwell GPUs. This will let developers build even larger neural networks and accelerate the bandwidth-intensive portions of deep learning training.
– Memory chips are stacked on top of one another and adjacent to the GPU, thus reducing the distance that data need to travel from memory to GPU and vice versa. Accelerates communication and improves power efficiency.
NVLink – for Faster Data Movement
– Data moves between GPUs and CPUs 5-12 times faster than compared to PCIe, great for deep learning.
– Allow twice as many GPUs in a system to work together in deep learning computations. In addition, CPUs and GPUs can connect in new ways to enable more flexibility and energy efficiency in server design compared to PCI-E.



