차세대 AI 가속기 및 GPU 클러스터
Next-Gen AI Accelerators & GPU Superclusters
텐서 코어, 고밀도 HBM 인터커넥트 및 거대 언어 모델 병렬 학습 아키텍처
연계 실시간 산업 동향
엔비디아, 차세대 블랙웰 Ultra B200 및 GB200 NVL72 랙스케일 수냉 슈퍼클러스터 공급 개시
130kW 랙 전력을 소화하는 GB200 NVL72 랙스케일 시스템이 주요 클라우드 서비스 사업자(CSP)에 납품되기 시작했다. 5세대 NVLink 1.8TB/s 양방향 인터커넥트로 초대형 LLM 학습 효율을 4배 끌어올렸다.
'AI 깐부동맹'과 'K수소 4대천왕'[광화문]
지난 28일(현지시간) 미국 뉴욕 맨해튼에서 열린 코리아소사이어티 연례 갈라(Gala)는 '찐친(진짜 친한 친구)'을 의미하는 '깐부'들의 동맹을 재확인하는 자리였다. 이날 젠슨 황 엔비디아 최고경영자(CEO)가 한미 우호 증진에 힘쓴 인사에게 수여하는 밴플리트상...
Apple’s A20 Solder Joint Diagram Leak Shows The Base iPhone 18 Will Arrive With Multiple Compromises, Such As Inferior Packaging And A Lower GPU Core Count
The A20 Pro will have several upgrades over the standard A20, as we’ve come to know from the latest solder joint diagram leak, which highlights a m...
엔비디아 투자 생태계 올라탄 韓 VC
Deep Dive 연계 학술 논문
Rack-Scale AI Computing with 130kW Two-Phase Immersion Cooling and 3.2 Tbps Co-Packaged Optics
D. Paterson, W. J. Dally, C. E. Kozyrakis et al.
Design and empirical thermal-electrical characterization of an integrated 72-GPU rack cluster operating at 130kW total power, utilizing dielectric fluorochemical two-phase immersion cooling with PUE of 1.025 and 3.2 Tbps CPO optical links.
ZTA-Q: an Open-source RISC-V Platform for Accurate Quantized CNN Inference
Yike Li, Ajay Kumar M, Vishnu PS et al.
Low-precision inference is widely adopted in edge AI to reduce computational cost and memory footprint. However, existing open-source accelerator platforms provide limited end-to-end support for CNNs following the standard TensorFlow Lite integer inference scheme. This paper presents ZTA-Q, an open-source RISC-V-based platform that enables accurate deployment of TensorFlow Lite INT8 models. In addition to extending operator support, ZTA-Q provides a configurable post-processing datapath for studying how circuit-level approximations, including reduced multiplier precision, shared shift scaling, and simplified rounding, affect model accuracy. The proposed system is implemented on a Digilent Arty A7-100T FPGA and operates at 83.3 MHz. Evaluations on representative CNN models show that with LUT, register, and DSP overheads of 26.3%, 12.6%, and 150%, respectively, ZTA-Q limits the degradation in both top-1 and top-5 accuracy to within 0.25 percentage points.
On the transmission of floating-point perturbations in flow-dependent filter-width formulations in Large-Eddy Simulation
Valerio D'Alessandro, Alessio Piccolo, Matteo Falone et al.
Heterogeneous high--performance computing architectures expose numerical algorithms to perturbations arising from the non--associativity of floating--point arithmetic. In Large--Eddy Simulation (LES), similar numerical effects may become relevant when they affect the filter--width entering the subgrid scale (SGS) model. This work investigates this mechanism for the least--squares (LSQ) based filter--width formulation, focusing on how floating--point effects are generated, transmitted, and coupled with the resolved flow. We show that, for fixed resolved kinematics, the LSQ filter--width is logarithmically non-expansive but not strictly contractive with respect to perturbations of the mesh metrics. Consequently, small disturbances may be transmitted with little attenuation through strongly directional filter-width responses. To mitigate this sensitivity, we introduce a scalar max--min compression of the directional mesh scales together with a bounded modulation based on the resolved velocity gradient. The resulting formulation reroutes floating--point perturbations through the filter-width operator, improving robustness while preserving the flow--dependent character of the original LSQ construction. The framework is assessed on heterogeneous CPU and GPU architectures for flow past a circular cylinder at Re=3900 and the Taylor--Green vortex at Re=1600. In the former, nearly one--to--one transmission of relative metric disturbances can become relevant when the transmitted perturbations interact with shear-layer transition. By contrast, on orthogonal Taylor--Green vortex grids, the accumulation pathway is structurally absent and no comparable macroscopic response develops. These results suggest floating--point sensitivity matters for LES filter--width formulations in heterogeneous computing environments.