2026
- Accepted [MICRO'26]SCOPE: Array-Native Nonlinear Acceleration for LLM Inference via Shape-Constrained Neural Approximation
- Accepted [MICRO'26]MFNNS: A Multiplier-Free Near-Memory Accelerator for Graph-based ANNS
- Accepted [MICRO'26]Ultra-DSP: A Universal Lossless DSP Overpacking Framework for Low-Bit LLM Inference
- Accepted [MICRO'26]Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators
- Accepted [MICRO'26]RidgeBridge: Random Access Optimized Interconnect Architecture for Scalable Graph Random Walks
- [ISCA'26]UniCore: A Bit-Width Scalable GEMM Unit for Unified LLM Inference
- [ICML'26]MixFP4: Extending NVFP4 to Mixed Micro-Format via Scale-Bit Reuse and Tensor Core Co-design
- [IPDPS'26]DSTREE: Data-Driven Synchronous Traversals for Decision Forests on GPUs
- [HPCA'26]RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAs
2025
- [MICRO'25]AxCore: A Quantization-Aware Approximate GEMM Unit for LLM Inference
- [MICRO'25]X-SET: An Efficient Graph Pattern Matching Accelerator With Order-Aware Parallel Intersection Units
- [ICCAD'25]OA-LAMA: An Outlier-Adaptive LLM Inference Accelerator with Memory-Aligned Mixed-Precision Group Quantization
- [TACO'25]Advancing Matrix Operations for High-Performance and Memory-Efficient Automata Processing on GPUs
- [LCTES'25]Graphitron: A Domain Specific Language for FPGA-Based Graph Processing Accelerator Generation
- [APNet'25]Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
- [DAC'25]April: Accuracy-Improved Floating-Point Approximation For Neural Network Accelerators
- [SIGMOD'25]Clementi: Efficient Load Balancing and Communication Overlap for Multi-FPGA Graph Processing
2023
- [SIGMOD'23]LightRW: FPGA Accelerated Graph Dynamic Random Walks
2022
- [MICRO'22]ReGraph: Scaling Graph Processing on HBM-enabled FPGAs with Heterogeneous Pipelines
- [TRETS'22]ThunderGP: Resource-efficient graph processing framework on FPGAs with HLS
2021
- [FPGA'21]ThunderGP: HLS-based graph processing framework on FPGAs
- [DAC'21]Skew-oblivious data routing for data intensive applications on FPGAs with HLS
- [SC'21]ThundeRiNG: generating multiple independent random number sequences on FPGAs
2020
- [CIDR'20]Is FPGA useful for hash joins?
- [VLDB'20]G3: when graph neural networks meet parallel graph processing systems on GPUs
2019
- [CIKM'19]Deploying hash tables on die-stacked high bandwidth memory
- [FPL'19]On-the-fly parallel data shuffling for graph processing on OpenCL-based FPGAs
- [JCST'19]A survey on graph processing accelerators: Challenges and opportunities
- [FPT'19]OBFS: OpenCL based BFS optimizations on software programmable FPGAs