Publications
publications by categories in reversed chronological order. generated by jekyll-scholar.
† marks equal authorship (authors contribute equally to the publication).
2026
-
CXL AnySSD: A composable CXL SSD using a CXL Type-2 device and any SSDsIEEE/ACM International Symposium on Microarchitecture (MICRO), Oct 2026
-
AMG: AMX-GPU cooperative acceleration of billion-scale ANNS via GEMM reformulationIEEE/ACM International Symposium on Microarchitecture (MICRO), Oct 2026
-
Elastic LLM serving on FPGAs via partial reconfiguration: a case study using the Altera FPGA AI SuiteIEEE/ACM International Symposium on Microarchitecture (MICRO), Oct 2026
-
NiF: hybrid in-flash and in-storage acceleration for approximate nearest neighbor searchIEEE/ACM International Symposium on Microarchitecture (MICRO), Oct 2026
-
ARC: adaptive reconfigurable CXL hotness monitoringIEEE/ACM International Symposium on Microarchitecture (MICRO), Oct 2026
-
Rethinking compression for CXL memory expanders at hyperscaleIEEE/ACM International Symposium on Microarchitecture (MICRO), Oct 2026*The alphabetical order
-
SOSPBorges: A low-latency distributed shared log on a CXL memory/SSD hybridACM Symposium on Operating Systems Principles (SOSP), Sep 2026
-
HyGIN: Hybrid CPU/GPU-initiated communication for mixture-of-experts trainingJuly–December 2026
-
TOCSTigon: A distributed database for a CXL podACM Transactions on Computer Systems (TOCS), to appear, 2026
-
MAC: Metadata acceleration for sustainable performance in big-data systems with CXL DRAMUSENIX Symposium on Operating Systems Design and Implementation (OSDI), Jul 2026
-
HCSXCENA MX1 CXL computational memory deviceIEEE Hot Chips Symposium (HCS) , Aug 2026
-
A Simulator for LLM inference systems exploiting CXL memory poolsJuly–December 2026
-
KiF: Accelerating low-batch LLM inference using in-flash KV cacheJuly–December 2026
-
PVAC: A RowHammer mitigation architecture exploiting per-victim-row countingIEEE/ACM International Symposium on Computer Architecture (ISCA) , Jun 2026
-
MMC: Metadata migration for efficient memory management in CXL DRAM systemsJuly–December 2026
-
Characterizing system-level trade-offs of Intel DSA for user-space IPC offloadingJuly–December 2026
-
Capacity-latency tradeoffs in CXL memory expander at hyperscaleJuly–December 2026
-
ISPASSCompiler and system optimizations for gem5 simulatorIEEE International Symposium on Performance Analysis of System and Software (ISPASS) , Apr 2026
-
-
-
MemSOS: OS-guided selective memory mirroringIEEE International Symposium on High-Performance Computer Architecture , Jan 2026
-
TCASCharacterizing the intrinsic bank-level accuracy vs. energy trade-off of SRAM-based analog in-memory computing architectures in 28 nm CMOSIEEE Transactions on Circuits and Systems I: Regular Papers , 2026
-
S&PSoK: Systematizing a decade of architectural RowHammer defenses through the lens of streaming algorithmsIEEE Symposium on Security and Privacy (S&P) , Mar 2026
-
2025
-
EPEPSTemperature-Dependent SPICE Models for UCIe InterconnectsIEEE Conference on Electrical Performance of Electronics Packaging and Systems , Oct 2025
-
-
-
DRAM fault classification through large-scale field monitoring for robust memory RAS management2025 58th IEEE/ACM International Symposium on Microarchitecture (MICRO) , Oct 2025
-
Stratum: system-hardware co-design with tiered monolithic 3D-DRAM for efficient MoE serving2025 58th IEEE/ACM International Symposium on Microarchitecture (MICRO) , Oct 2025
-
CABANA: Cluster-aware query batching for accelerating billion-scale ANNS with Intel AMXDec 2025
-
HINT: A hardware platform for intra-host NIC traffic and SmartNIC emulationDec 2025
-
Time series machine learning models for precise SSD access latency predictionDec 2025
-
-
srNAND: A novel NAND flash organization for Enhanced Small Read Throughput in SSDs2025
-
Dynamic Load Balancer in Intel® Xeon® Scalable Processor: Performance analyses, enhancements, and guidelinesIEEE/ACM International Symposium on Computer Architecture (ISCA) , 2025
-
A4: Microarchitecture-aware LLC management for datacenter servers with emerging I/O devicesIEEE/ACM International Symposium on Computer Architecture (ISCA) , 2025
-
LIA: A single-GPU LLM inference acceleration with cooperative AMX-enabled CPU-GPU computation and CXL offloadingIEEE/ACM International Symposium on Computer Architecture (ISCA) , 2025
-
Universal predicate pushdown to smart storageIEEE/ACM International Symposium on Computer Architecture (ISCA) , 2025
-
Hybrid SLC-MLC RRAM mixed-signal processing-in-memory architecture for Transformer acceleration via gradient redistributionIEEE/ACM International Symposium on Computer Architecture (ISCA) , 2025
-
ISPASSIntel In-Memory Analytics Accelerator: performance characterization and guidelinesIEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2025
-
DACA full-system, programmable, and extensible in-memory computing simulation framework for deep learningIEEE/ACM Design Automation Conference (DAC) , 2025
-
TVLSIExploiting chiplet integration technology for fast high-capacity DRAM modulesIEEE Transactions on Very Large Scale Integration Systems (TVLSI) , 2025
-
Hardware-Accelerated Kernel-Space Memory Compression Using Intel QATIEEE Computer Architecture Letters , 2025
-
X-PPR: Post Package Repair for CXL MemoryIEEE Computer Architecture Letters , 2025
-
Warped-Compaction: Maximizing GPU Register File Bandwidth Utilization via Operand Compaction2025 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2025
-
Cooperative Memory Deduplication with Intel Data Streaming AcceleratorIEEE Computer Architecture Letters , 2025
2024
-
AttAcc! Unleashing the power of PIM for batched transformer-based generative model inference29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2024
-
An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , 2024
-
A quantitative analysis and guidelines of data streaming accelerator in modern intel xeon scalable processors29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2024
-
Tandem processor: Grappling with emerging operators in neural networks29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2024
-
Exploiting Intel® Advanced Matrix Extensions (AMX) for Large Language Model InferenceIEEE Computer Architecture Letters , 2024
-
DRAMScope: Uncovering DRAM Microarchitecture and Characteristics by Issuing Memory Commands2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) , 2024
-
Intel accelerators ecosystem: An soc-oriented perspective: Industry product2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) , 2024
-
TAROT: A CXL SmartNIC-Based Defense Against Multi-bit Errors by Row-Hammer AttacksProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) , Apr 2024
-
Spade: Sparse pillar-based 3d object detection accelerator for autonomous driving2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , Feb 2024
-
Transforming the Hybrid Cloud for Emerging AI WorkloadsarXiv preprint arXiv:2411.13239 , 2024
-
-
Hal: Hardware-assisted load balancing for energy-efficient snic-host cooperative computing2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) , 2024
-
Demystifying a CXL Type-2 Device: A Heterogeneous Cooperative Computing Perspective2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2024
-
Computer Architecture Having Selectable Parallel and Serial Communication Channels Between Processors and Memory2024
-
Yield-Aware Interposer Design for UCIe Interconnects2024 IEEE 33rd Conference on Electrical Performance of Electronic Packaging and Systems (EPEPS) , 2024
-
Tandem Processor: Grappling with Emerging Operators in Neural Networks2024
2023
-
Unleashing the potential of pim: Accelerating large batched inference of transformer-based generative modelsIEEE Computer Architecture Letters , 2023
-
Making sense of using a smartnic to reduce datacenter tax from slo and tco perspectives2023 IEEE International Symposium on Workload Characterization (IISWC) , 2023
-
ISPASSAnalyzing Energy Efficiency of a Server with a SmartNIC under SLO Constraints2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2023
-
Defensive ml: Defending architectural side-channels with adversarial obfuscationarXiv preprint arXiv:2302.01474 , 2023
2022
-
Pipegcn: Efficient full-graph training of graph convolutional networks with pipelined feature communicationarXiv preprint arXiv:2203.10428 , 2022
-
Unlocking the power of inline Floating-Point operations on programmable switches19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) , 2022
-
An FPGA-based RNN-T Inference Accelerator with PIM-HBM2022
-
-
Coordinated Science Laboratory 70th Anniversary Symposium: The Future of ComputingarXiv preprint arXiv:2210.08974 , 2022
-
HAMS: Hardware Automated Memory-over-Storage for Large-scale Memory Expansion13rd Annual Non-Volatile Memories Workshop (NVMW), 2022 , 2022
2021
-
Hardware architecture and software stack for PIM based on commercial DRAM technology: Industrial product2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA) , 2021
-
25.4 a 20nm 6gb function-in-memory dram based on hbm2 with a 1.2 tflops programmable computing unit using bank-level parallelism for machine learning applications2021 IEEE International Solid-State Circuits Conference (ISSCC) , 2021
-
Network-centric architecture and algorithms to accelerate distributed training of neural networks2021
-
Diag: a dataflow-inspired architecture for general-purpose processors26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2021
-
QEI: Query acceleration can be generic and efficient in the cloud2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , 2021
-
Greendimm: Os-assisted dram power management for dram with a sub-array granularity power-down state2021
-
In-Memory Near-Data Approximate Acceleration2021
-
FlatFlash system for byte granularity accessibility of memory in a unified memory-storage hierarchy2021
2020
-
A 16-GB 640-GB/s HBM2E DRAM with a data-bus window extension technique and a synergetic on-die ECC schemeIEEE Journal of Solid-State Circuits , 2020
-
Babelfish: Fusing address translations for containers2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) , 2020
-
Freac cache: Folded-logic reconfigurable computing in the last level cache2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2020
-
BDS-GCN: Efficient full-graph training of graph convolutional nets with partition-parallelism and boundary sampling2020
-
Graphic processor unit providing reduced storage costs for similar operands2020
2019
2018
-
A network-centric hardware/algorithm co-design to accelerate distributed training of deep neural networks2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2018
-
Application-transparent near-memory processing architecture with memory channel network2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2018
-
Leveraging power-performance relationship of energy-efficient modern DRAM devicesIEEE Access , 2018
2017
-
Collaborative (cpu+ gpu) algorithms for triangle counting and truss decomposition on the minsky architecture: Static graph challenge: Subgraph isomorphism2017 IEEE High Performance Extreme Computing Conference (HPEC) , 2017
-
Ncap: Network-driven packet context-aware power management for client-server architecture2017 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2017
2016
-
Write-after-read hazard prevention in GPGPUSIMWorkshop on Deplicating, Deconstructing, and Debunking (WDDD) , 2016
2015
2014
2013
-
IMPROVING MEMORY RELIABILITY POWER AND PERFORMANCE USING MIXED-CELL DESIGNS.Intel Technology Journal , 2013
2012
-
Parameter variation at near threshold voltage: The power efficiency versus resilience tradeoffUniversity of Illinois, Tech. Rep , 2012
2011
2010
-
Combating aging with the colt duty cycle equalizer2010 43rd Annual IEEE/ACM International Symposium on Microarchitecture , 2010
-
Optimal algorithm for profile-based power gating: A compiler technique for reducing leakage on execution units in microprocessors2010 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) , 2010
2009
-
Method and apparatus improving performance of a digital memory array device2009
2008
2007
2006
2005
2004
2003
-
Power analyzer for pocket computing (papc)University of Michigan, Tech. Rep. , 2003
2002
2001
-
VLSI Implementation of Binaural Spatializer using FIR Head-Related Transfer Function (HRTF)2001