AI Computing Overtakes GPU Power: Apple A20 Pro Reveals Dual 16-Core NPU Focus on On-Device AI Models

Deep News
1 hour ago

Recent reports indicate that the Apple A20 Pro chip has already set a new single-core performance record in Geekbench 6 for smartphone processors, with multi-core results showing gains of up to 27% compared to the A19 Pro and M3, and even surpassing the M5 Max. Now, a fresh detail has surfaced regarding this system-on-chip: its dual 16-core neural engine can deliver peak computing power that exceeds the performance of the built-in 7-core GPU, but only when the workload is specially tailored for the NPU.

This information comes from a well-known tipster, Max Weinbach, who shared the insight on the X platform. He highlighted an intriguing fact about the A20 Pro and M6: as long as tasks are specifically optimized for the neural engine, the dual 16-core configuration can provide a performance boost that outpaces the GPU. However, if the neural engine is tasked with graphics-intensive workloads or specific animations, it operates so slowly that it becomes nearly unusable, meaning such jobs must still be handled by the GPU.

This architectural shift suggests that the A20 Pro's design philosophy is now tilted toward AI computing power. When running on-device AI models at the 3 billion (3B) parameter scale or handling the prompt pre-fill stage of quantized models, the dual 16-core NPU delivers exceptional execution efficiency. In other words, the NPU's pipeline has been specifically optimized for AI prompt processing, allowing it to efficiently take over many tasks that previously required GPU involvement, thereby freeing up more system resources.

Beyond model processing, the A20 Pro's NPU also takes on a wide range of vision and audio-specific tasks. For instance, real-time camera depth estimation, background blur, text recognition, and image segmentation are all handled by the NPU at the underlying level. On the audio side, real-time speech recognition and ambient sound enhancement now benefit from hardware-level acceleration, with the NPU running these tasks at a much higher energy efficiency ratio than traditional cores.

That said, the NPU is not a universal solution. When dealing with non-quantized 32-bit (FP32) or 64-bit floating-point math, as well as tasks involving dynamic tensor shapes, the system still falls back to the CPU or GPU for computation. Additionally, for the single-token generation phase of large language models, the NPU's scheduling overhead makes the CPU and GPU more suitable for such tasks.

Currently, the first devices equipped with the A20 Pro, including the iPhone 18 Pro, iPhone 18 Pro Max, and iPhone Duo, have not yet shipped. As these devices gradually hit the market, more tasks specifically designed for the dual 16-core neural engine are expected to emerge, which will make it clearer just how much practical benefit this upgrade truly delivers.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10