DOI,IST REx ID,Research Group,Title of publication
10.1007/978-3-031-85747-8_6,21257,"DaAl,GradSch",Sparse Fine-Tuning for Inference Acceleration of Large Language Models
10.1145/3710848.3710871,19877,DaAl,MARLIN: Mixed-precision auto-regressive parallel inference on Large Language Models
null,18975,DaAl,Error feedback can accurately compress preconditioners
null,18977,DaAl,SpQR: A sparse-quantized representation for near-lossless LLM weight compression
10.5281/ZENODO.14213091,19884,DaAl,MARLIN: Mixed-precision auto-regressive parallel inference on Large Language Models
null,17456,DaAl,L-GreCo: Layerwise-adaptive gradient compression for efficient data-parallel deep learning
null,18113,"DaAl,GradSch",Extreme compression of large language models via additive quantization
null,18121,DaAl,SPADE: Sparsity-guided debugging for deep neural networks
10.15479/at:ista:17485,17485,"GradSch,DaAl","Compressing large neural networks: Algorithms, systems and scaling laws"
null,18062,DaAl,Scaling laws for sparsely-connected foundation models
null,18061,DaAl,QMoE: Sub-1-bit compression of trillion parameter models
null,17378,DaAl,OPTQ: Accurate post-training quantization for generative pre-trained transformers
null,14458,DaAl,SparseGPT: Massive language models can be accurately pruned in one-shot
null,17059,DaAl,SPDY: Accurate pruning with speedup guarantees
10.18653/v1/2022.emnlp-main.279,17088,DaAl,The optimal BERT surgeon: Scalable and accurate second-order pruning for large language models
null,17087,DaAl,Optimal brain compression: A framework for accurate post-training quantization and pruning
null,11463,DaAl,M-FAC: Efficient matrix-free approximations of second-order information
null,8724,"DaAl,ChLa",On the sample complexity of adversarial multi-source PAC learning
