High-Bandwidth NoC Design for AI Accelerator Platforms

Authors

  • VenkataKrishna Brahmam Pogadadanda Senior RTL Design Engineer at AMD-Xilinx, USA. Author

DOI:

https://doi.org/10.63282/3117-5481/AIJCST-V7I2P111

Keywords:

Network-On-Chip (Noc), AI Accelerators, High-bandwidth Communication, On-chip Interconnect, Latency Optimization, Congestion Control, ASIC Design, Scalable Architectures

Abstract

The rapid development of artificial intelligence (AI), machine learning, and data-centric computing has significantly amplified the demand for highly efficient on-chip communication topologies in today's accelerator platforms. AI accelerators feature several processing parts, memory hierarchies, tensor engines and special computational units that need to move large amounts of data with low latency and optimal efficiency. Traditional bus-based communication systems cannot provide the bandwidth and scalability requirements of modern AI workloads, particularly deep learning training and inference with large neural network models. This leads to the necessity of Network-on-Chip (NoC) architectures as a vital solution for scalable and high-performance interconnectivity in AI accelerator systems. High-bandwidth NoCs provide efficient packet-based communication across distributed processing nodes, leading to enhanced parallelism, better resource utilization, and improved overall system throughput. The design of an effective NoC for AI accelerator platforms has become an important issue of scalability, communication latency, congestion control, power efficiency, and dependability. More processing cores and memory interfaces make low-latency communication harder to preserve and network traffic congestion harder to avoid. Further, interconnect fabrics with high bandwidths have large power consumption and may limit the temperature for the progress of semiconductor technology. However, the use of NoC in large-scale AI systems is hindered by reliability difficulties such as packet loss, deadlocks, fault tolerance, and traffic imbalance. This paper presents a complete methodology for the design of high-bandwidth Network-on-Chip (NoC) for AI accelerator platforms, focusing on scalable topology selection, adaptive routing algorithms, congestion-aware traffic management, low-power communication techniques, and methods for enhancing dependability. We demonstrate a case study of a practical AI accelerator to explain implementation considerations and performance trade-offs for real AI workloads. Experimental research shows that the proposed methodology may significantly optimize bandwidth utilization, reduce communication latency, alleviate congestion hotspots and enhance energy efficiency compared to typical NoC systems.

References

[1] Jain, Vikram, and Marian Verhelst. "Networks-on-chip to enable large-scale multi-core ml acceleration." Towards Heterogeneous Multi-core Systems-on-Chip for Edge Machine Learning: Journey from Single-core Acceleration to Multi-core Heterogeneous Systems. Cham: Springer Nature Switzerland, 2023. 143-161.

[2] Bhowmik, Biswajit R. "Ai technology in networks-on-chip." Industrial Transformation. CRC Press, 2022. 99-128.

[3] Suryadevara, S. S. K., & Nakirikanti, S. (2024). Blockchain-Backed Content Authenticity Verification Framework. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(1), 242-252. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I1P125

[4] Gaddam, R. R., & Krishna, K. (2023). KFP v2 Artifact-Centric ML Pipeline Governance. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 4(2), 142-153. https://doi.org/10.63282/3050-9262.IJAIDSML-V4I2P116

[5] Nabavinejad, Seyed Morteza, et al. "An overview of efficient interconnection networks for deep neural network accelerators." IEEE Journal on Emerging and Selected Topics in Circuits and Systems 10.3 (2020): 268-282.

[6] Vppalapati, M. (2024). Cooling Domains as First-Class Failure Boundaries in Storage Architecture. American International Journal of Computer Science and Technology, 6(2), 96-106. https://doi.org/10.63282/3117-5481/AIJCST-V6I2P110

[7] Muppaneni, K. (2024). Progressive Web Apps: Offline UX Benchmarking. International Journal of Emerging Trends in Computer Science and Information Technology, 5(2), 174-183. https://doi.org/10.63282/3050-9246.IJETCSIT-V5I2P119

[8] Yang, Lei, et al. "Co-exploring neural architecture and network-on-chip design for real-time artificial intelligence." 2020 25th Asia and South Pacific Design Automation Conference (ASP-DAC). IEEE, 2020.

[9] Shiramalla, R. (2024). Secure Multi-Cloud API Orchestration between Salesforce, Oracle CPQ, and Azure. American International Journal of Computer Science and Technology, 6(3), 102-113. https://doi.org/10.63282/3117-5481/AIJCST-V6I3P108

[10] Srigadde, B. R., & Devaraju, J. M. (2024). Building a Reusable AI Connection Utility Class. International Journal of Emerging Research in Engineering and Technology, 5(2), 188-200. https://doi.org/10.63282/3050-922X.IJERET-V5I2P119

[11] Jain, Vikram, et al. "PATRONoC: Parallel AXI transport reducing overhead for networks-on-chip targeting multi-accelerator DNN platforms at the edge." 2023 60th ACM/IEEE Design Automation Conference (DAC). IEEE, 2023.

[12] Muppaneni, R. K. (2023). AI-Driven Forecasting in Dynamics 365 Sales: What Businesses Need to Know. International Journal of AI, BigData, Computational and Management Studies, 4(1), 168-176. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V4I1P117

[13] Katangoori, S., & Katangoori, A. (2021). AI-Augmented Data Governance: Enabling Intelligent Access, Lineage, and Compliance across Hybrid Clouds. American International Journal of Computer Science and Technology, 3(6), 36-45. https://doi.org/10.63282/3117-5481/AIJCST-V3I6P104

[14] Rella¹, Bhanu Prakash Reddy, et al. "AI-Engine-Based Acceleration for High-Performance Programmable System-on-Chip Designs." Journal of Computational Analysis and Applications 32.1 (2024).

[15] Gaddam, R. R. (2023). Progressive Delivery for Models with Quality KPIs. American International Journal of Computer Science and Technology, 5(4), 33-47. https://doi.org/10.63282/3117-5481/AIJCST-V5I4P104

[16] Muppaneni, K., & Palem, V. (2024). Micro-Frontend Design Patterns for Multi-Framework Applications. International Journal of Emerging Research in Engineering and Technology, 5(3), 181-190. https://doi.org/10.63282/3050-922X.IJERET-V5I3P120

[17] Chen, Yiran, et al. "A survey of accelerator architectures for deep neural networks." Engineering 6.3 (2020): 264-274.

[18] Vanapalli, Satyendra Kumar. "Business Process Reengineering Using CRM Workflow Automation." International Journal of Intelligent Automation & Robotics Engineering 7.2 (2024): 01-17.

[19] Suryadevara, S. S. K. (2024). Resilient Multi-CDN Delivery Model Using AI-Based Traffic Switching for Global AEM Deployments. International Journal of Emerging Trends in Computer Science and Information Technology, 5(3), 191-200. https://doi.org/10.63282/3050-9246.IJETCSIT-V5I3P119

[20] Boutros, Andrew, Eriko Nurvitadhi, and Vaughn Betz. "Architecture and application co-design for beyond-FPGA reconfigurable acceleration devices." IEEE Access 10 (2022): 95067-95082.

[21] Takkalapally, D., & Takkellapally, M. R. (2024). AI-SynPerf: Synthetic Data Intelligence Framework for 5G Mobile Performance Simulation. International Journal of Emerging Trends in Computer Science and Information Technology, 5(1), 182-194. https://doi.org/10.63282/3050-9246.IJETCSIT-V5I1P118

[22] Kommuru, Madhurima, and Appala Nooka Kumar Doodala. "Performance Bottlenecks Necks in Data Heavy Python Applications." International Journal of Applied Data Science & Modern Computing 6.2 (2023): 01-19.

[23] Parakala, A. (2023). Citizen-Facing Automation: Chatbots and Self-Service in Public Services. International Journal of AI, BigData, Computational and Management Studies, 4(4), 108-118. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V4I4P112

[24] Li, Shenggao, et al. "High-bandwidth chiplet interconnects for advanced packaging technologies in AI/ML applications: Challenges and solutions." IEEE Open Journal of the Solid-State Circuits Society 4 (2024): 351-364.

[25] Shiramalla, R. (2023). Optimizing Cross-Platform Enterprise Integrations Using Workato: A Case Study of Salesforce and Oracle SaaS Applications. International Journal of Emerging Trends in Computer Science and Information Technology, 4(1), 232-243. https://doi.org/10.63282/3050-9246.IJETCSIT-V4I1P124

[26] Vppalapati, M. (2024). Power-Bound Storage Design: Architecting Systems for Electrical Scarcity. International Journal of AI, BigData, Computational and Management Studies, 5(1), 208-217. https://doi.org/10.63282/3050-9416.IJAIBDCMS-V5I1P121

[27] Park, Seunghyun, and Daejin Park. "Low-Power Scalable TSPI: A Modular Off-Chip Network for Edge AI Accelerators." IEEE Access 12 (2024): 141448-141459.

[28] Allenki, S. S. (2024). Building Scalable Data Replication Pipelines for Real-Time Analytics. American International Journal of Computer Science and Technology, 6(1), 71-81. https://doi.org/10.63282/3117-5481/AIJCST-V6I1P108

[29] Kumar Doodala, A. N. (2024). Service Virtualization for API-First development: A shift-Left Testing Strategy. American International Journal of Computer Science and Technology, 6(4), 50-58. https://doi.org/10.63282/3117-5481/AIJCST-V6I4P105

[30] Raha, Arnab, et al. "Design considerations for edge neural network accelerators: An industry perspective." 2021 34th International Conference on VLSI Design and 2021 20th International Conference on Embedded Systems (VLSID). IEEE, 2021.

[31] Takkalapally, D. (2024). ShiftLeft-AI: Machine Learning Framework for Proactive Performance Assurance in CI/CD Pipelines. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(4), 285-296. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I4P126

[32] Muppaneni, R. K. (2023). Low-Code Revolution: How Power Platform Extends Dynamics 365 Capabilities. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 4(3), 162-171. https://doi.org/10.63282/3050-9262.IJAIDSML-Vssss4I3P119

[33] Akhoon, Mohd Saqib, et al. "High performance accelerators for deep neural networks: A review." Expert Systems 39.1 (2022): e12831.

[34] Kommuru, M. (2024). Assessing the Limitations of LangChain in Production Environments. American International Journal of Computer Science and Technology, 6(4), 71-83. https://doi.org/10.63282/3117-5481/AIJCST-V6I4P107

[35] Parakala, A. (2023). Vendor Highlights – IoT, AI, and Process Mining. International Journal of Emerging Trends in Computer Science and Information Technology, 4(4), 135-146. https://doi.org/10.63282/3050-9246.IJETCSIT-V4I4P115

[36] Hu, Xianghong, et al. "High-performance reconfigurable DNN accelerator on a bandwidth-limited embedded system." ACM Transactions on Embedded Computing Systems 22.6 (2023): 1-20.

[37] Kumar Doodala, A. N. (2024). Validating UX consistency Across Omnichannel Platform. American International Journal of Computer Science and Technology, 6(6), 87-97. https://doi.org/10.63282/3117-5481/AIJCST-V6I6P109

[38] Mandal, Sumit K., Anish Krishnakumar, and Umit Y. Ogras. "Energy-efficient networks-on-chip architectures: Design and run-time optimization." Network-on-Chip Security and Privacy (2021): 55-75.

[39] Vanapalli, Satyendra Kumar. "Identifying Operational Inefficiencies through CRM Data Analysis." International Journal of Data Engineering and Intelligent Computing 7.2 (2024): 01-13.

[40] Allenki, S. S. (2024). Automating Backups and Recovery: Reducing Manual Work by Over 50%. International Journal of Emerging Research in Engineering and Technology, 5(1), 166-176. https://doi.org/10.63282/3050-922X.IJERET-V5I1P119

[41] Srigadde, B. R. (2024). Agents, LLMs, and Salesforce with Multi-Cloud Provider (MCP). International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(3), 277-288. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I3P127

[42] Asadikouhanjani, Mohammadreza, and Seok-Bum Ko. "Enhancing the utilization of processing elements in spatial deep neural network accelerators." IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 40.9 (2020): 1947-1951.

[43] Taluri, R. (2024). A Hybrid Data Engineering and Generative AI Architecture for Intelligent Data Governance, Metadata Management, and Automated Data Quality Assessment on AWS. International Journal of Artificial Intelligence, Data Science, and Machine Learning, 5(2), 230-240. https://doi.org/10.63282/3050-9262.IJAIDSML-V5I2P126

Downloads

Published

2025-03-30

Issue

Section

Articles

How to Cite

[1]
V. B. Pogadadanda, “High-Bandwidth NoC Design for AI Accelerator Platforms”, AIJCST, vol. 7, no. 2, pp. 134–146, Mar. 2025, doi: 10.63282/3117-5481/AIJCST-V7I2P111.

Similar Articles

31-40 of 263

You may also start an advanced similarity search for this article.