Skip to content
Squarehuang's DejaVu
扉页
Initializing search
SeeleVolle/SeeleVolle.github.io
简介
课程笔记
论文阅读
技术积累
个人闲谈
无用收藏
Squarehuang's DejaVu
SeeleVolle/SeeleVolle.github.io
简介
简介
课程笔记
课程笔记
None
大数据存储与计算技术
论文阅读
论文阅读
Transformer Review
Wanda:A Simple and Effective Pruning Approach for Large Language Models
SparseGPT:Massive Language Models Can Be Accurately Pruned in One-Shot
PowerInfer:Fast Large Language Model Serving with a Consumer-grade GPU
Sheared LLAMA:Accelerating Language Model Pre-training via Structured Pruning
De jaVu:Conditional Regenerative Learning to Enhance Dense Prediction:
ATOM:LOW-BIT QUANTIZATION FOR EFFICIENT AND ACCURATE LLM SERVING:
GEAR:An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM:
PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-Optimization
LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table Lookup
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
FastDecode:High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
DEFT: FLASH TREE-ATTENTION WITH IO-AWARENESS FOR EFFICIENT TREE-SEARCH-BASED LLM INFERENCE
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
技术积累
技术积累
常用工具
常用工具
Mysql Notes
Docker Notes
SSH Notes
Git Notes
PackageManager Notes
FileSystem Notes
CMake 学习笔记
VIM Notes
编程相关 | 代码分析 | Linux
编程相关 | 代码分析 | Linux
GoF23 Notes
modernC++
CUDA
PIM Learning
Linux Learning
SwiftTransformer
Machine Learning
Machine Learning
Linear Algebra
ML Basic Knowledge
其他知识
其他知识
一些奇怪的知识合集
稀疏矩阵
BT
内存硬件基础知识
个人闲谈
个人闲谈
Squarehuang's 归档
Squarehuang's 归档
2024
2023
2022
Squarehuang's 分类
Squarehuang's 分类
动漫
学习
日常
无用收藏
无用收藏
站点收藏
Markdown语法
看番记录
扉页
¶
一些工具性的文档....
“少小知名翰墨场,十年心事只凄凉。”
Back to top