SOTAVerified

CPU

Papers

Showing 51–100 of 2231 papers

TitleStatusHype
Benchmarking of CPU-intensive Stream Data Processing in The Edge Computing Systems—0
Bang for the Buck: Vector Search on Cloud CPUs—0
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference—0
Fast Differentiable Modal Simulation of Non-linear Strings, Membranes, and PlatesCode1
FloE: On-the-Fly MoE Inference on Memory-constrained GPU—0
Plexus: Taming Billion-edge Graphs with 3D Parallel GNN Training—0
Edge-GPU Based Face Tracking for Face Detection and Recognition Acceleration—0
Supporting renewable energy planning and operation with data-driven high-resolution ensemble weather forecast—0
The Unreasonable Effectiveness of Discrete-Time Gaussian Process Mixtures for Robot Policy Learning—0
RetroInfer: A Vector-Storage Approach for Scalable Long-Context LLM Inference—0
Morello: Compiling Fast Neural Networks with Dynamic Programming and Spatial CompressionCode1
Spill The Beans: Exploiting CPU Cache Side-Channels to Leak Tokens from Large Language Models—0
GPRat: Gaussian Process Regression with Asynchronous TasksCode0
Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning—0
Accelerated 3D-3D rigid registration of echocardiographic images obtained from apical window using particle filter—0
Mesh-Learner: Texturing Mesh with Spherical HarmonicsCode1
GPU accelerated program synthesis: Enumerate semantics, not syntax!—0
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration—0
Dynamic Superblock Pruning for Fast Learned Sparse Retrieval—0
Blockchain Meets Adaptive Honeypots: A Trust-Aware Approach to Next-Gen IoT Security—0
ThyroidEffi 1.0: A Cost-Effective System for High-Performance Multi-Class Thyroid Carcinoma Classification—0
MetaDSE: A Few-shot Meta-learning Framework for Cross-workload CPU Design Space Exploration—0
NNTile: a machine learning framework capable of training extremely large GPT language models on a single node—0
Chinese-Vicuna: A Chinese Instruction-following Llama-based ModelCode7
BitNet b1.58 2B4T Technical Report—0
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures—0
MULTI-LF: A Unified Continuous Learning Framework for Real-Time DDoS Detection in Multi-Environment Networks—0
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length FloatCode4
Understanding and Optimizing Multi-Stage AI Inference Pipelines—0
OVERLORD: Ultimate Scaling of DataLoader for Multi-Source Large Foundation Model Training—0
aweSOM: a CPU/GPU-accelerated Self-organizing Map and Statistically Combined Ensemble Framework for Machine-learning Clustering Analysis—0
Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention—0
Wavefront Estimation From a Single Measurement: Uniqueness and Algorithms—0
Towards On-Device Learning and Reconfigurable Hardware Implementation for Encoded Single-Photon Signal Processing—0
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints—0
WoundAmbit: Bridging State-of-the-Art Semantic Segmentation and Real-World Wound Care—0
GPU-accelerated Evolutionary Many-objective Optimization Using Tensorized NSGA-IIICode3
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE InferenceCode2
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home ClustersCode0
IAEmu: Learning Galaxy Intrinsic Alignment CorrelationsCode0
Accurate GPU Memory Prediction for Deep Learning Jobs through Dynamic Analysis—0
Exploring energy consumption of AI frameworks on a 64-core RV64 Server CPU—0
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism—0
Accelerating IoV Intrusion Detection: Benchmarking GPU-Accelerated vs CPU-Based ML Libraries—0
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching—0
SCRec: A Scalable Computational Storage System with Statistical Sharding and Tensor-train Decomposition for Recommendation Models—0
Solving the Best Subset Selection Problem via Suboptimal AlgorithmsCode0
GPU-centric Communication Schemes for HPC and ML Applications—0
Deep Learning Model Deployment in Multiple Cloud Providers: an Exploratory Study Using Low Computing Power Environments—0
Concorde: Fast and Accurate CPU Performance Modeling with Compositional Analytical-ML Fusion—0
Show:102550
← PrevPage 2 of 45Next →

No leaderboard results yet.