More
More
Stars
Pytorch implementation for FAT: learning low-bitwidth parametric representation via frequency-aware transformation
[NeurIPS-2024] 📈 Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies https://arxiv.org/abs/2407.13623