PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization-Reference-Cited by-同舟云学术

PTQ4ViT: Post-training Quantization for Vision Transformers with Twin Uniform Quantization

Published:2022 Issue: Volume: Page:191-207
ISSN:0302-9743
Container-title:Lecture Notes in Computer Science
language:
Short-container-title:

Author:

Yuan Zhihang,Xue Chenhao,Chen Yiqi,Wu Qiang,Sun Guangyu

Publisher

Springer Nature Switzerland

Link

https://link.springer.com/content/pdf/10.1007/978-3-031-19775-8_12

Reference28 articles.

1. Banner, R., Nahshan, Y., Soudry, D.: Post training 4-bit quantization of convolutional networks for rapid-deployment. In: Wallach, H.M., Larochelle, H., Beygelzimer, A., d’Alché-Buc, F., Fox, E.B., Garnett, R. (eds.) Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, 8–14 December 2019. Vancouver, BC, Canada, pp. 7948–7956 (2019). https://proceedings.neurips.cc/paper/2019/hash/c0a62e133894cdce435bcb4a5df1db2d-Abstract.html

2. Brown, T.B., et al.: Language models are few-shot learners. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, 6–12 December 2020. virtual (2020). https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html

3. Lecture Notes in Computer Science;N Carion,2020

4. Choi, J., Wang, Z., Venkataramani, S., Chuang, P.I., Srinivasan, V., Gopalakrishnan, K.: PACT: parameterized clipping activation for quantized neural networks. CoRR abs/1805.06085 (2018). http://arxiv.org/abs/1805.06085

5. Choukroun, Y., Kravchik, E., Yang, F., Kisilev, P.: Low-bit quantization of neural networks for efficient inference. In: 2019 IEEE/CVF International Conference on Computer Vision Workshops, ICCV Workshops 2019, Seoul, Korea (South), 27–28 October 2019, pp. 3009–3018. IEEE (2019). https://doi.org/10.1109/ICCVW.2019.00363

Cited by 30 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Minimize Quantization Output Error with Bias Compensation;CAAI Artificial Intelligence Research;2025-07

2. A survey of FPGA and ASIC designs for transformer inference acceleration and optimization;Journal of Systems Architecture;2024-10

3. Computer Vision Model Compression Techniques for Embedded Systems:A Survey;Computers & Graphics;2024-10

4. VitBit: Enhancing Embedded GPU Performance for AI Workloads through Register Operand Packing;Proceedings of the 53rd International Conference on Parallel Processing;2024-08-12

5. Q-A2NN: Quantized All-Adder Neural Networks for Onboard Remote Sensing Scene Classification;Remote Sensing;2024-06-30