Lightning-AI/lightning-thunder
View on GitHubquantization: process tensors on meta device directly, maybe implement CPU quantization (if it is easy)
Open
#1,111 opened on Sep 6, 2024
good first issuetransforms
Repository metrics
- Stars
- (1,460 stars)
- PR merge metrics
- (PR metrics pending)
Description
Currently the BitsAndBytesLinearQuant4bit for submodule always calls bitsandbytes.functional.quantize_4bit. This is somewhat touchy for CPU tensors because quantize_4bit only works on GPU tensors but it is outright not so nice for meta tensors, where we only would need to get the right shapes.