DP
DeepTech
목차: Ch A1.1 차세대 AI 가속기 및 GPU 클러스터
Ch A1.1난이도 L3기술 온톨로지 정의

차세대 AI 가속기 및 GPU 클러스터

Next-Gen AI Accelerators & GPU Superclusters

정본 핸드북 원문

텐서 코어, 고밀도 HBM 인터커넥트 및 거대 언어 모델 병렬 학습 아키텍처

관련 개체:nvidiaamdgooglerebellyonsfuriosa
연계 개념:hbm
📖
본 개념은 AI 하드웨어 (AI Hardware) 전략 온톨로지 정본으로 등록되어 실시간 글로벌 인텔리전스 및 최신 학술 논문(ArXiv) 스트림과 연동 중입니다.

연계 실시간 산업 동향

총 4건
Tom's Hardware10. 1.

엔비디아, 차세대 블랙웰 Ultra B200 및 GB200 NVL72 랙스케일 수냉 슈퍼클러스터 공급 개시

130kW 랙 전력을 소화하는 GB200 NVL72 랙스케일 시스템이 주요 클라우드 서비스 사업자(CSP)에 납품되기 시작했다. 5세대 NVLink 1.8TB/s 양방향 인터커넥트로 초대형 LLM 학습 효율을 4배 끌어올렸다.

블랙웰 B200/GB200 AI 가속기원문
머니투데이10. 1.

'AI 깐부동맹'과 'K수소 4대천왕'[광화문]

지난 28일(현지시간) 미국 뉴욕 맨해튼에서 열린 코리아소사이어티 연례 갈라(Gala)는 '찐친(진짜 친한 친구)'을 의미하는 '깐부'들의 동맹을 재확인하는 자리였다. 이날 젠슨 황 엔비디아 최고경영자(CEO)가 한미 우호 증진에 힘쓴 인사에게 수여하는 밴플리트상...

'엔비디아' 일치 · Ch A1.1 대규모 AI 학습 GPU원문
Wccftech10. 1.

Apple’s A20 Solder Joint Diagram Leak Shows The Base iPhone 18 Will Arrive With Multiple Compromises, Such As Inferior Packaging And A Lower GPU Core Count

The A20 Pro will have several upgrades over the standard A20, as we’ve come to know from the latest solder joint diagram leak, which highlights a m...

'gpu' 일치 · Ch A1.1 대규모 AI 학습 GPU원문
'엔비디아' 일치 · Ch A1.1 대규모 AI 학습 GPU원문

Deep Dive 연계 학술 논문

총 3편
arXiv:2601.077212026. 1.

Rack-Scale AI Computing with 130kW Two-Phase Immersion Cooling and 3.2 Tbps Co-Packaged Optics

D. Paterson, W. J. Dally, C. E. Kozyrakis et al.

Design and empirical thermal-electrical characterization of an integrated 72-GPU rack cluster operating at 130kW total power, utilizing dielectric fluorochemical two-phase immersion cooling with PUE of 1.025 and 3.2 Tbps CPO optical links.

도메인: AI 하드웨어 시리즈PDF 원문 열람
arXiv:2610.01867v12026. 10.

ZTA-Q: an Open-source RISC-V Platform for Accurate Quantized CNN Inference

Yike Li, Ajay Kumar M, Vishnu PS et al.

Low-precision inference is widely adopted in edge AI to reduce computational cost and memory footprint. However, existing open-source accelerator platforms provide limited end-to-end support for CNNs following the standard TensorFlow Lite integer inference scheme. This paper presents ZTA-Q, an open-source RISC-V-based platform that enables accurate deployment of TensorFlow Lite INT8 models. In addition to extending operator support, ZTA-Q provides a configurable post-processing datapath for studying how circuit-level approximations, including reduced multiplier precision, shared shift scaling, and simplified rounding, affect model accuracy. The proposed system is implemented on a Digilent Arty A7-100T FPGA and operates at 83.3 MHz. Evaluations on representative CNN models show that with LUT, register, and DSP overheads of 26.3%, 12.6%, and 150%, respectively, ZTA-Q limits the degradation in both top-1 and top-5 accuracy to within 0.25 percentage points.

도메인: AI 하드웨어 시리즈PDF 원문 열람
arXiv:2610.01743v12026. 10.

On the transmission of floating-point perturbations in flow-dependent filter-width formulations in Large-Eddy Simulation

Valerio D'Alessandro, Alessio Piccolo, Matteo Falone et al.

Heterogeneous high--performance computing architectures expose numerical algorithms to perturbations arising from the non--associativity of floating--point arithmetic. In Large--Eddy Simulation (LES), similar numerical effects may become relevant when they affect the filter--width entering the subgrid scale (SGS) model. This work investigates this mechanism for the least--squares (LSQ) based filter--width formulation, focusing on how floating--point effects are generated, transmitted, and coupled with the resolved flow. We show that, for fixed resolved kinematics, the LSQ filter--width is logarithmically non-expansive but not strictly contractive with respect to perturbations of the mesh metrics. Consequently, small disturbances may be transmitted with little attenuation through strongly directional filter-width responses. To mitigate this sensitivity, we introduce a scalar max--min compression of the directional mesh scales together with a bounded modulation based on the resolved velocity gradient. The resulting formulation reroutes floating--point perturbations through the filter-width operator, improving robustness while preserving the flow--dependent character of the original LSQ construction. The framework is assessed on heterogeneous CPU and GPU architectures for flow past a circular cylinder at Re=3900 and the Taylor--Green vortex at Re=1600. In the former, nearly one--to--one transmission of relative metric disturbances can become relevant when the transmitted perturbations interact with shear-layer transition. By contrast, on orthogonal Taylor--Green vortex grids, the accumulation pathway is structurally absent and no comparable macroscopic response develops. These results suggest floating--point sensitivity matters for LES filter--width formulations in heterogeneous computing environments.

도메인: AI 하드웨어 시리즈PDF 원문 열람