랙스케일 인프라 및 전력망 연계 시스템
Rack-Scale AI Computing & Power Grid
NVL72 등 단일 거대 컴퓨터 통합 스케일아웃 네트워킹과 기가와트 전력 인프라
연계 실시간 산업 동향
※ 해당 개념의 24시간 내 직접 속보가 없어 AI 하드웨어 시리즈 대표 실시간 공급망 뉴스를 연동합니다.
엔비디아, 차세대 블랙웰 Ultra B200 및 GB200 NVL72 랙스케일 수냉 슈퍼클러스터 공급 개시
130kW 랙 전력을 소화하는 GB200 NVL72 랙스케일 시스템이 주요 클라우드 서비스 사업자(CSP)에 납품되기 시작했다. 5세대 NVLink 1.8TB/s 양방향 인터커넥트로 초대형 LLM 학습 효율을 4배 끌어올렸다.
구글, 우주 데이터센터 첫 궤도 실험…TPU 실은 위성 발사 성공
(서울=연합뉴스) 정주호 기자 = 구글의 인공지능(AI) 칩을 실은 시제품 위성이 스페이스X 로켓에 실려 궤도에 오르면서 빅테크의 '우주 데이터...
브로드컴, 102.4Tbps 스위치용 공동패키징 광학(CPO) 엔진 공개… 실리콘 포토닉스 인터커넥트 시대
전기 신호 인터커넥트의 거리 및 전력 한계를 해결하는 1.6T 실리콘 포토닉스 광학 엔진이 스위치 패키지 위에 직접 통합됐다. AI 클러스터 광학 네트워킹 전력 소모를 50% 절감한다.
'AI 깐부동맹'과 'K수소 4대천왕'[광화문]
지난 28일(현지시간) 미국 뉴욕 맨해튼에서 열린 코리아소사이어티 연례 갈라(Gala)는 '찐친(진짜 친한 친구)'을 의미하는 '깐부'들의 동맹을 재확인하는 자리였다. 이날 젠슨 황 엔비디아 최고경영자(CEO)가 한미 우호 증진에 힘쓴 인사에게 수여하는 밴플리트상...
Murata Unveils CPO Road Map Integrating Optical, Electrical and Thermal Technologies
Murata Manufacturing Co. has unveiled a technology road map for co-packaged optics (CPO) for artificial intelligence data centers, detailing the de...
Deep Dive 연계 학술 논문
MoE-CORE: Coordinated Expert Offloading and Residency for Memory-Constrained MoE Inference
Ke Yang, Yongji Gao, Xushi Li et al.
Sparse expert activation reduces MoE models' computation, yet expert weights can exceed limited device memory. Offloading makes inference feasible on a compact AI appliance but exposes host-to-device transfers to the inference path. We present MoE-CORE, a system that coordinates expert offloading and residency for memory-constrained MoE inference. It stages complete expert layers in alternating buffers during prefill. During decode, it combines nonuniform layer-wise cache capacity, domain-informed initialization, routing-history-aware replacement, and cross-layer prefetching. The main configuration executes router-selected experts exactly; an optional score-based substitution path handles eligible low-score misses. The main comparison uses 1K- and 128-token output caps for MoE-CORE and vLLM Prefetch, respectively. Across five workloads per model, MoE-CORE records a mean time per output token (TPOT) of 38.0-44.8 ms versus 1268.9-1269.1 ms for the evaluated vLLM Prefetch configuration on DeepSeek-V4-Flash-W4A8; the corresponding values on GLM-5.2-W4A8C8 are 206.6-220.5 and 5941.5-5941.8 ms. Under an 84-GB NPU-memory cap, the best measured DeepSeek GSM8K configuration achieves a TPOT of 21.5 ms with approximate expert substitution and multi-token prediction (MTP) at depth 2. These results support coordinated expert residency and transfer scheduling under a device-memory constraint. The code is here.
Timing-Driven Logic Remapping with Local Physical Context
Zijian Jiang, Hongyang Pan, Cunqing Lan et al.
The timing behavior of a mapped circuit depends on both its logic implementation and the physical environment in which that implementation is realized. Revisiting mapping decisions after placement therefore requires a search procedure that accounts for surrounding timing constraints, fanout loads, and interconnect effects. We study local remapping in this setting and develop a framework that couples discrete mapping search with physical implementation feedback. Timing-critical regions are isolated through bounded windows whose interfaces retain the context of the surrounding circuit. Within each window, a mixed-integer formulation jointly selects logic cuts, signal polarities, and library cells under a delay model informed by estimated locations and interconnect parasitics. A continuous relaxation filters the search space before discrete optimization produces alternative implementations with similar modeled timing and different structural choices. These implementations are reconstructed and assessed through legalization, routing-based parasitic estimation, and timing analysis. Physically validated improvements are incorporated into the design, and the updated context guides subsequent searches. The framework provides a systematic way to revisit local logic implementations while accounting for their interaction with an existing placement.
ZTA-Q: an Open-source RISC-V Platform for Accurate Quantized CNN Inference
Yike Li, Ajay Kumar M, Vishnu PS et al.
Low-precision inference is widely adopted in edge AI to reduce computational cost and memory footprint. However, existing open-source accelerator platforms provide limited end-to-end support for CNNs following the standard TensorFlow Lite integer inference scheme. This paper presents ZTA-Q, an open-source RISC-V-based platform that enables accurate deployment of TensorFlow Lite INT8 models. In addition to extending operator support, ZTA-Q provides a configurable post-processing datapath for studying how circuit-level approximations, including reduced multiplier precision, shared shift scaling, and simplified rounding, affect model accuracy. The proposed system is implemented on a Digilent Arty A7-100T FPGA and operates at 83.3 MHz. Evaluations on representative CNN models show that with LUT, register, and DSP overheads of 26.3%, 12.6%, and 150%, respectively, ZTA-Q limits the degradation in both top-1 and top-5 accuracy to within 0.25 percentage points.
Open-Source Multi-Wire SPI Readout for Wearable Ultrasound Probes
Federico Villani, Soumyo Bhattacharjee, Lisa Odermatt et al.
Wearable ultrasound probes must transfer increasingly large acquisition payloads while maintaining compact, low-power electronics. In TinyProbe, the current bottleneck in data transfer occurs between the acquisition FPGA and the wireless system controller. This work presents an open-source, multi-wire SPI readout interface that uses serial command and address phases followed by a build-time-selectable dual- or quad-lane payload phase that is intended to address this bottleneck by increasing the potential bandwidth over the wifi limit while retaining compatibility with the Microcontroller-centric wearable US architecture. The interface emulates a serial flash memory, enabling compatibility with a broad range of microcontroller families and their existing peripheral interfaces. On the FPGA, the data path connects the existing acquisition FIFOs to the SPI interface through clock-domain crossing, sample reshaping, and packing into 32-bit words. Dual-SPI readout is integrated into the existing IGLOO2/SiWG917 TinyProbe architecture and verified at an SCLK frequency of 5 MHz. A separate Kria K26 testbed is used to characterize the FPGA SPI interface independently of the acquisition and wireless subsystems, demonstrating error-free transfers at SCLK frequencies up to 66 MHz. These measurements identify the SiWG917 multi-lane SPI implementation as the next bandwidth-limiting component and motivate a future upgrade of the system controller. The HDL and MCU implementations are released under a permissive open-source license.
U-Sonic: An Open-Source 8-Channel Ultrasound Transmit IP in a 130 nm RISC-V SoC
Federico Villani, Nico Canzani, Marc-André Wessner et al.
Miniaturized ultrasound (US) probes require programmable and synchronized transmit (TX) excitation across multiple elements, while existing compact platforms often rely on limited microcontroller (MCU) pulse generators or closed-source fixed-function pulser devices. We present U-Sonic, an open-source digital US TX peripheral integrated into a 32-bit RISC-V system-on-chip (SoC). The implemented SoC integrates 8 pulser cores, while the parameterized architecture supports up to 16 channels. Each core generates single- or dual-tone bursts with programmable period, duty cycle, pulse count, polarity, and idle level, together with optional inverted stop pulses for active damping. A shared memory-mapped Open Bus Interface (OBI) enables synchronous start and stop of arbitrary channel subsets and supports composite bipolar, gated, and three-level excitation schemes. Functional correctness was verified in Verilator against a Python golden model over 4379 checked cycles across directed and randomized configurations, and confirmed on a Terasic DE10-Lite field-programmable gate array (FPGA). The design was synthesized and placed-and-routed in IHP 130 nm. The post-layout area in kilo gate equivalents (kGE), scales as 1.65 kGE plus 1.66 kGE per channel. The 8-channel instance occupies 14.9 kGE, corresponding to approximately 14.3% of the 104 kGE SoC. The register-transfer level (RTL), register descriptions, verification collateral, and software support are released as open source.