-
Notifications
You must be signed in to change notification settings - Fork 34.1k
Insights: huggingface/transformers
July 22, 2026 – July 29, 2026
Overview
Summary
Excluding merges, 63 authors have pushed 70 commits to main and 289 commits to all branches.
On main, 644 files have changed and there have been 17,443 additions and 13,703 deletions
70 pull requests merged by 33 people
- [docs] response_template when serving
#47626 merged
1 hour agoJul 30, 2026 - [docs] MTP support
#47301 merged
1 hour agoJul 30, 2026 - Fix GPT-2 c_proj depth scaling initialization
#47459 merged
2 hours agoJul 30, 2026 - [docs] MPS graph cache
#47304 merged
2 hours agoJul 30, 2026 - Fix fp8_linear compilability
#47623 merged
2 hours agoJul 30, 2026 - byebye torch 2.4
#47609 merged
4 hours agoJul 29, 2026 - [`Kernels`] Refactor function handling
#46883 merged
4 hours agoJul 29, 2026 - Compressed tensors fp8
#47216 merged
4 hours agoJul 29, 2026 - Vectorize NoRepeatNGramLogitsProcessor and remove its host sync
#47571 merged
4 hours agoJul 29, 2026 - Fix CUDA Graph breaking host to device copy from scalar tensor allocation
#47547 merged
6 hours agoJul 29, 2026 - Fix some processors
#47608 merged
7 hours agoJul 29, 2026 - skip fsdp tests when backend is mps
#47601 merged
9 hours agoJul 29, 2026 - CI: Add serge review relay workflow and review rules
#47610 merged
11 hours agoJul 29, 2026 - [docs] Exporters
#47374 merged
14 hours agoJul 29, 2026 - Fix model parallel device mismatch in `create_bidirectional_sliding_window_mask`
#47560 merged
yesterdayJul 29, 2026 - CI: use a single function for GH calls
#47474 merged
yesterdayJul 28, 2026 - Kernels and loaders robustification
#47334 merged
yesterdayJul 28, 2026 - Fix A.X-K2 fp8 modules_to_not_convert normalization for the gated-norm MLP
#47578 merged
yesterdayJul 28, 2026 - Remove redundant guarding for distributed
#47570 merged
2 days agoJul 28, 2026 - Add FSDP plans to all models
#47165 merged
2 days agoJul 28, 2026 - Fix modular for mamba packages
#47494 merged
2 days agoJul 28, 2026 - Remove deprecated conversion in Kimi
#47581 merged
2 days agoJul 28, 2026 - CI: use transformers-ci daily workflow with OTEL
#47360 merged
2 days agoJul 27, 2026 - Fix slow tensor path in _check_special_mm_tokens
#47580 merged
2 days agoJul 27, 2026 - CI: let's run integration failure cron at 10pm
#47537 merged
2 days agoJul 27, 2026 - Better and more extensive tests for RoPE
#46912 merged
2 days agoJul 27, 2026 - Fix mamba2 family decode and simplify all reshape ops
#47569 merged
2 days agoJul 27, 2026 - Speed up image preprocessing for vision-language models
#47453 merged
2 days agoJul 27, 2026 - [DiffusionGemma] Cast the decoder padding mask to bool
#47295 merged
2 days agoJul 27, 2026 - Fix NemotronH: Register `"mlp"` in the cache layer-type mappings
#47535 merged
2 days agoJul 27, 2026 - Fix `config.per_layer_config[layer_type]` lookup speed
#47539 merged
2 days agoJul 27, 2026 - General maintenance
#47517 merged
2 days agoJul 27, 2026 - :rotating_light: Processors update the rest
#46556 merged
2 days agoJul 27, 2026 - fixed the benchmark script with DistributedConfig
#47568 merged
2 days agoJul 27, 2026 - Fix TP inference for tied embedding
#47503 merged
3 days agoJul 27, 2026 - Add AXK2 from SKT
#47528 merged
5 days agoJul 25, 2026 - Fix `causal_conv1d_fn` positional `activation` colliding with hub kernel's `seq_idx`
#47527 merged
5 days agoJul 25, 2026 - Fix failing tests for mimo_v2_flash
#47284 merged
5 days agoJul 25, 2026 - Refactor all linear attention models to latest best standards for convolution
#47452 merged
5 days agoJul 25, 2026 - Make tokenization_mistral_common importable without mistral_common installed
#47397 merged
5 days agoJul 25, 2026 - Use --flake-runs=1 in check_bad_commit.py for PR comment CI
#47522 merged
5 days agoJul 25, 2026 - Deprecate the old response_schema
#47320 merged
5 days agoJul 24, 2026 - fix: add pickle support to _LazyConfigMapping for spawn multiprocessing
#46026 merged
5 days agoJul 24, 2026 - Fix potential ReDoS by escaping tokenizer filename used as regex pattern
#47498 merged
5 days agoJul 24, 2026 - Fix vision position-embedding init width fallback in Phi4Multimodal
#47509 merged
5 days agoJul 24, 2026 - Delete old deprecations
#47518 merged
5 days agoJul 24, 2026 - Run only @slow tests in PR comment CI
#47521 merged
5 days agoJul 24, 2026 - Fix incorrect type hint
#47519 merged
5 days agoJul 24, 2026 - Deprecate CB config in gen configuration
#47291 merged
last weekJul 24, 2026 - Allow metal-flash-sdpa for OpenAIPrivacyFilter on MPS
#46740 merged
last weekJul 24, 2026 - [Offloading] [Bugfix] Fix fully offloaded model saving
#47336 merged
last weekJul 24, 2026 - CI: fix torchaudio pinning +proper break in rnnt
#47422 merged
last weekJul 24, 2026 - add_axk1
#46867 merged
last weekJul 24, 2026 - Fix Qwen2.5-Omni Token2Wav DiT rotary embedding layout (interleaved cos/sin)
#47403 merged
last weekJul 24, 2026 - CI: add reproduce mode to serge verify caller
#47492 merged
last weekJul 24, 2026 - feat[vLLM x v5]: Make audio optional and support multiple audios in VibeVoice ASR processor
#47483 merged
last weekJul 23, 2026 - [fix][whisper]: fix max_new_tokens handling
#46795 merged
last weekJul 23, 2026 - Isolate MLA KV expansion to make it easier to bypass
#47460 merged
last weekJul 23, 2026 - Update serve chat parsing
#46267 merged
last weekJul 23, 2026 - Fix loss alignment and Trainer token counting for encoder decoder models
#46903 merged
last weekJul 23, 2026 - Fix Gemma4 audio feature dtype mismatch in masked_scatter
#47482 merged
last weekJul 23, 2026 - Fix the HunyuanVL's torchvision backend
#47499 merged
last weekJul 23, 2026 - Use new `per_layer_config` for Gemma 4 so that heterogeneous attention config is explicit
#47384 merged
last weekJul 23, 2026 - [DiffusionGemma] Support gradient checkpointing
#46572 merged
last weekJul 23, 2026 - Fix failing tests for zaya
#47268 merged
last weekJul 23, 2026 - tipsv2_dpt: fix failing tests for XPU
#47292 merged
last weekJul 23, 2026 - Fix some failed test cases related with XPU Expectations
#47173 merged
last weekJul 23, 2026 - add paged attention tests support for XPU
#47163 merged
last weekJul 23, 2026 - Add FSDP CI and end-to-end FSDP tests + save fsdp
#47357 merged
last weekJul 23, 2026 - fix: guard DTensor import in sharding_utils.py for PyTorch < 2.5
#47481 merged
last weekJul 23, 2026
41 pull requests opened by 28 people
- [Bugfix] [Modeling] Fix DSV4 norm
#47486 opened
last weekJul 23, 2026 - Fix CodeLlama tokenizer dropping leading whitespace on decode
#47488 opened
last weekJul 23, 2026 - skip invalid test cases for inkling tests
#47493 opened
last weekJul 23, 2026 - Acc fix in xpu
#47500 opened
last weekJul 23, 2026 - update Dockerfile for xpu torch2.13
#47502 opened
last weekJul 23, 2026 - [Mistral] Add native tekken tokenizer support to AutoTokenizer
#47507 opened
last weekJul 24, 2026 - [Quantization] Fix CT `apply_quantization_config`
#47510 opened
last weekJul 24, 2026 - Raise a clear error when a token is both forced and suppressed
#47511 opened
last weekJul 24, 2026 - Allow position_ids_start=2 on DataCollatorWithFlattening for RoBERTa etc.
#47525 opened
5 days agoJul 25, 2026 - [Chat Parsing] Type inline tool-call arguments from the calling tool's JSON Schema
#47529 opened
5 days agoJul 25, 2026 - QA: mlinter rule 20
#47536 opened
4 days agoJul 25, 2026 - Hoist special-token lookups in wav2vec2 decode paths and drop a dead filter in wav2vec2_phoneme
#47557 opened
3 days agoJul 26, 2026 - Fix Pix2StructTextAttention init using hidden_size instead of d_kv
#47558 opened
3 days agoJul 26, 2026 - [serge] Fix 2 integration tests for model `got_ocr2` failing with `other` (other (2))
#47564 opened
3 days agoJul 27, 2026 - [serge] Fix 2 integration tests for model `vivit` failing with `output_mismatch` (tensor values differ (2))
#47566 opened
3 days agoJul 27, 2026 - Modularize qwen-format vision processors
#47573 opened
2 days agoJul 27, 2026 - [nit] use requires_backends
#47576 opened
2 days agoJul 27, 2026 - TP dtensor inference
#47579 opened
2 days agoJul 27, 2026 - Bump the actions group across 1 directory with 14 updates
#47584 opened
2 days agoJul 27, 2026 - Support solar-open2 model
#47585 opened
2 days agoJul 27, 2026 - fix npu check
#47587 opened
2 days agoJul 28, 2026 - feat[vLLM]: Make VoxtralProcessor compatible with the Transformers modeling backend for vLLM
#47588 opened
2 days agoJul 28, 2026 - [Experimental / Do not Review] Add multi-method export + QNN (Qualcomm HTP) backend to the ExecuTorch exporter
#47594 opened
2 days agoJul 28, 2026 - Fix dtype casting when dequantizing
#47596 opened
2 days agoJul 28, 2026 - 🚨 Simplify all Rotary modules
#47598 opened
2 days agoJul 28, 2026 - Potential fix for code scanning alert no. 265: Environment variable built from user-controlled sources
#47600 opened
yesterdayJul 28, 2026 - Align OlmoHybrid to use a native cache
#47604 opened
yesterdayJul 28, 2026 - [serge] Fix 4 integration tests for model `hunyuan_vl` failing with `other` (other (4))
#47606 opened
yesterdayJul 28, 2026 - [serge] Fix 2 integration tests for model `granite4_vision` failing with `other` (other (2))
#47607 opened
yesterdayJul 28, 2026 - Resolve the Hub revision once per load instead of passing a private _commit_hash around
#47611 opened
yesterdayJul 29, 2026 - Don't warn about class defaults when diffing a config
#47613 opened
yesterdayJul 29, 2026 - feat[vLLM]: Add return_text_replacement_offsets to TextKwargs
#47614 opened
yesterdayJul 29, 2026 - Gate eager distributed imports on torch.distributed.is_available()
#47615 opened
yesterdayJul 29, 2026 - [serge] Fix 2 integration tests for model `regnet` failing with `output_mismatch` (tensor values differ (2))
#47616 opened
19 hours agoJul 29, 2026 - better guarding to handle torch compiled with USE_DISTRIBUTED=0
#47619 opened
13 hours agoJul 29, 2026 - Always tie structural weights, regardless of `tie_word_embeddings`
#47620 opened
12 hours agoJul 29, 2026 - Refresh MiniMax context window defaults to match official checkpoints
#47621 opened
9 hours agoJul 29, 2026 - Drop multimodal inputs natively in prepare_inputs_for_generation if not in prefill
#47622 opened
9 hours agoJul 29, 2026 - [Improvement] Make gated delta rule more explicit and support per-channel decay
#47625 opened
8 hours agoJul 29, 2026 - Fix out-of-vocab default bos/eos token ids in Siglip text configs
#47627 opened
41 minutes agoJul 30, 2026 - fix processor config nested key fallback
#47628 opened
10 minutes agoJul 30, 2026
78 issues closed by 13 people
- Adding kNN language modeling and Machine Translation
#8346 closed
2 hours agoJul 30, 2026 - [Bug] GPT2 c_proj weights incorrectly initialised
#47458 closed
2 hours agoJul 30, 2026 - "ValueError: Tokenizer class TokenizersBackend does not exist or is not currently imported"
#47553 closed
2 hours agoJul 30, 2026 - Hybrid models
#36646 closed
4 hours agoJul 29, 2026 - Marian RNN conversion support
#36651 closed
4 hours agoJul 29, 2026 - Speed up image processors - cast to array before BatchFeature
#31205 closed
15 hours agoJul 29, 2026 - Running gpt-neo 2.7B with less than 13GB of system memory like Colab
#11276 closed
16 hours agoJul 29, 2026 - Supporting truncation from both ends of the sequence in BertTokenizerFast
#10082 closed
yesterdayJul 29, 2026 - Add thinking-budget support (max_thinking_tokens) for reasoning-capable chat models
#42111 closed
yesterdayJul 29, 2026 - Adding native support to load GGUF models using transformers
#38063 closed
yesterdayJul 29, 2026 - Batch Decoding of LMs will cause different outputs with different batch size
#25921 closed
yesterdayJul 29, 2026 - [serge] integration failure triage - 2026-07-27
#47567 closed
yesterdayJul 28, 2026 - device_map loading OOMs on conversion-op transients: WeightConverter contiguous() runs on destination GPU, unbudgeted by max_memory packing
#47404 closed
yesterdayJul 28, 2026 - Please implement DUMA: Reading Comprehension with Transposition Thinking
#10944 closed
2 days agoJul 28, 2026 - Add Flash Attention 2.0 for T5 Family
#27441 closed
2 days agoJul 28, 2026 - Request to add FLMR
#29103 closed
2 days agoJul 28, 2026 - Support symbolic tracing for NeoX models
#24937 closed
2 days agoJul 28, 2026 - Unified freezing interface
#13399 closed
2 days agoJul 28, 2026 - Feature Request: To add nested hierarchy retrieval from Donut response
#24722 closed
2 days agoJul 28, 2026 - Schedule Free Optimizers (pytorch) and Sophia optimizer
#30359 closed
2 days agoJul 27, 2026 - `KeyError: 'mlp'` when building a cache for Nemotron-H — `"mlp"` missing from the layer-type mappings
#47534 closed
2 days agoJul 27, 2026 - [`Layer Types`] Split keys across MLP and Attentions
#46245 closed
2 days agoJul 27, 2026 - [Bug] pytest-xdist workers race on captured_info.txt in patched testing utils
#45561 closed
2 days agoJul 27, 2026 - [serge] integration failure triage - 2026-07-26
#47563 closed
3 days agoJul 27, 2026 - Expand text-generation pipeline support for other causal models e.g., BigBirdForCausalLM
#12439 closed
3 days agoJul 27, 2026 - GQA Llama 13B slower than Llama 13B without GQA
#28425 closed
3 days agoJul 27, 2026 - [i18n-ur] Translating docs to Urdu (اردو)
#33139 closed
3 days agoJul 27, 2026 - example code problem
#27092 closed
3 days agoJul 27, 2026 - [testing] making network tests more reliable
#12061 closed
3 days agoJul 27, 2026 - Having a function to verify if checkpoint is valid
#31283 closed
3 days agoJul 27, 2026 - Add Evolla model
#36231 closed
3 days agoJul 27, 2026 - static cache implementation is not compatible with attn_implementation==flash_attention_2
#32040 closed
3 days agoJul 27, 2026 - replace roberta embedding with bge_base
#25824 closed
3 days agoJul 27, 2026 - Add ability to specify input device for ffmpeg_microphone()
#31074 closed
3 days agoJul 27, 2026 - model type `deepseek_v32`
#42590 closed
3 days agoJul 27, 2026 - Changes required to `save_model` for certain models (e.g., Phi 3.5 Vision)
#34690 closed
3 days agoJul 27, 2026 - Integrate Liger (Linkedin GPU Efficient Runtime) Kernel to HuggingFace
#32861 closed
3 days agoJul 27, 2026 - LLaMa-VID: An Image is Worth 2 Tokens in LLMs
#27980 closed
3 days agoJul 27, 2026 - Swapping `tqdm` to `rich`
#29064 closed
3 days agoJul 27, 2026 - Add support for causal language modeling for `DistilBertModel`
#34933 closed
3 days agoJul 27, 2026 - A model or a config like 'transformer_iwslt_de_en' for machine translation
#12386 closed
3 days agoJul 27, 2026 - [serge] integration failure triage - 2026-07-26
#47562 closed
3 days agoJul 27, 2026 - [serge] integration failure triage - 2026-07-25
#47538 closed
3 days agoJul 27, 2026 - Consider reverting CLIP refactor
#46644 closed
3 days agoJul 26, 2026 - Swin2SR image processor over-pads images when dimensions are already divisible by size_divisor
#46697 closed
3 days agoJul 26, 2026 - Several models hardcode `torch.float64` in their forward pass, crashing on MPS (Apple Silicon)
#46723 closed
3 days agoJul 26, 2026 - [serge] integration failure triage - 2026-07-24
#47515 closed
4 days agoJul 25, 2026 - Getting equivalent results between Transformer's resize and tf.image.resize
#27601 closed
5 days agoJul 25, 2026 - Abnormally High GPU Memory Consumption with OPT 350M Model Leading to OOM
#25419 closed
5 days agoJul 25, 2026 - Distil BART for text simplification
#11594 closed
5 days agoJul 25, 2026 - MeZo Forward Pass Implementation
#24264 closed
5 days agoJul 25, 2026 - `causal_conv1d_fn() got multiple values for argument 'seq_idx'` with the hub kernel
#47526 closed
5 days agoJul 25, 2026 - BLIVA
#26629 closed
5 days agoJul 25, 2026 - DeepSpeed Support Stage 3
#29254 closed
5 days agoJul 24, 2026 - Routing Transformers / Add Google PG-19 Models
#11686 closed
5 days agoJul 24, 2026 - [Qwen2VL] Image processor backend refactor (#43514) introduces pixel_values precision drift, causing bbox coordinate shift
#46688 closed
5 days agoJul 24, 2026 - [serge] integration failure triage - 2026-07-23
#47512 closed
5 days agoJul 24, 2026 - GenerationConfig serializes ContinuousBatchingConfig as null
#47039 closed
last weekJul 24, 2026 - [serge] integration failure triage - 2026-07-22
#47490 closed
last weekJul 24, 2026 - How to sending my request about parameters in inference API?
#21403 closed
last weekJul 24, 2026 - [Offloading] Cannot save disk-offloaded `Qwen/Qwen3-30B-A3B`
#47333 closed
last weekJul 24, 2026 - Add data2vec 2.0
#30805 closed
last weekJul 24, 2026 - Qwen2.5-Omni Token2Wav DiT uses mismatched RoPE layout (half-split cos/sin with interleaved rotate) — regression from #39847
#47328 closed
last weekJul 24, 2026 - Safety Checking Infrastructure for Text Generation
#41740 closed
last weekJul 24, 2026 - When using the Liger Kernel, torch.nn.functional.cross_entropy is called
#43039 closed
last weekJul 23, 2026 - PPFormulaNet training-loss double-shift: trains against labels[..., 1:] instead of labels
#46901 closed
last weekJul 23, 2026 - overflow_to_sample_mapping missing in in documentation
#9059 closed
last weekJul 23, 2026 - DTensor import in sharding_utils.py breaks on PyTorch < 2.5
#47480 closed
last weekJul 23, 2026 - NVIDIA RADIO-L
#40398 closed
last weekJul 23, 2026 - Porting Compressive Transformer to Huggingface
#15533 closed
last weekJul 23, 2026 - Support DBRX Model
#29911 closed
last weekJul 23, 2026 - [activations] pytorch-1.11+ Tanh Gelu Approximation
#15397 closed
last weekJul 23, 2026 - [serge] integration failure triage - 2026-07-21
#47465 closed
last weekJul 23, 2026 - `DTensor` import in `sharding_utils.py` breaks on PyTorch < 2.5
#47479 closed
last weekJul 23, 2026 - `CheckpointError` with PEFT + DeepSpeed ZeRO-3 + gradient checkpointing
#47254 closed
last weekJul 23, 2026 - BertTokenizer and BertTokenizerFast have different behavior when requested "return_overflowing_tokens"
#28900 closed
last weekJul 23, 2026
15 issues opened by 15 people
- Add Talkie (talkie-1930-13b)
#47618 opened
18 hours agoJul 29, 2026 - BOS and EOS token ID not within vocabulary warning for siglip2-base-patch16-naflex
#47612 opened
yesterdayJul 29, 2026 - [serge] integration failure triage - 2026-07-28
#47605 opened
yesterdayJul 28, 2026 - Importing `AutoImageProcessor` fails when PyTorch is built without `torch.distributed`
#47603 opened
yesterdayJul 28, 2026 - CodeLlamaTokenizer FIM/infilling drops correct leading whitespace on decode
#47602 opened
yesterdayJul 28, 2026 - Repetition penalty is gauge-dependent. Proposing opt-in normalized variant.
#47595 opened
2 days agoJul 28, 2026 - NemotronH: no longer respected after #47452 (unconditionally fetches Hub kernel)
#47577 opened
2 days agoJul 27, 2026 - Wav2Vec2 tokenizer rebuilds special-token lists per token in decode paths; phoneme variant carries a filter that never matches
#47556 opened
3 days agoJul 26, 2026 - Mamba2-family `cuda_kernels_forward` decode step is broken since #47452: missing seq-dim squeeze
#47532 opened
5 days agoJul 25, 2026 - Phi3/Phimoe: prepare_inputs_for_generation forwards logits_to_keep=None into forward() -> 4-dim logits -> torch.multinomial crash
#47530 opened
5 days agoJul 25, 2026 - Transformers 5.13 CPU memory growth loading Qwen3-235B-A22B with DeepSpeed ZeRO-3; 4.51.0 remains bounded
#47514 opened
last weekJul 24, 2026 - MistralCommonBackend API documentation appears to render the dummy implementation instead of the actual class
#47504 opened
last weekJul 23, 2026 - `DataCollatorWithFlattening` uses 0-based `position_ids`, silently degrading RoBERTa-family models
#47496 opened
last weekJul 23, 2026 - [BUG] CANINE: dynamic ONNX export fixes sequence length to example input
#47491 opened
last weekJul 23, 2026 - CodeLlama tokenizer drops leading whitespace on decode round-trip
#47487 opened
last weekJul 23, 2026
75 Unresolved conversations
Sometimes conversations happen on old items that aren't yet closed. Here is a list of all the Issues and Pull Requests with unresolved conversations.
- Port ESMC and ESMFold2 to Transformers
#46419 commented on
1 hour agoJul 30, 2026 • new comments - Add Unlimited OCR
#46836 commented on
9 minutes agoJul 30, 2026 • new comments - Support Granite Speech NAR (NLE)
#46031 commented on
1 hour agoJul 30, 2026 • new comments - wip step 3 7 flash
#46658 commented on
yesterdayJul 29, 2026 • new comments - Add support for Nemotron Omni
#46509 commented on
5 days agoJul 25, 2026 • new comments - [`FA`] Native torch integration
#45153 commented on
1 hour agoJul 30, 2026 • new comments - Executorch exporter fixes
#47243 commented on
5 hours agoJul 29, 2026 • new comments - Add MOSS-TTS-v1.5 (delay) and MOSS-Audio-Tokenizer
#46447 commented on
41 minutes agoJul 30, 2026 • new comments - Exportable kimi
#47096 commented on
1 hour agoJul 30, 2026 • new comments - DO NOT MERGE - model creation skill
#45213 commented on
11 hours agoJul 29, 2026 • new comments - 🚨 Enable SDPA (and other attention backends) for T5 and propagate to the T5 family
#47014 commented on
1 hour agoJul 30, 2026 • new comments - Add Fun-ASR-Nano model
#46180 commented on
2 hours agoJul 30, 2026 • new comments - [CB] [Major] Add multimodal support to continous batching
#46376 commented on
last weekJul 23, 2026 • new comments - Add Granite-swa and Granitemoe-swa model support
#47179 commented on
7 minutes agoJul 30, 2026 • new comments - Add pose estimation keypoint preprocessing to Sapiens2ImageProcessor
#47199 commented on
2 days agoJul 27, 2026 • new comments - Add V-JEPA 2.1 inference support
#45497 commented on
4 days agoJul 26, 2026 • new comments - model: Add NVIDIA Canary-1B-v2 to Transformers
#46825 commented on
2 hours agoJul 30, 2026 • new comments - add MTP for GLM OCR and other MROPEs
#47462 commented on
10 hours agoJul 29, 2026 • new comments - [Improvement] Formalize KV sharing
#47290 commented on
2 days agoJul 27, 2026 • new comments - Fix batched Qwen2.5/3-Omni audio generation
#47186 commented on
2 days agoJul 28, 2026 • new comments - Implement VibeVoice
#40546 commented on
2 days agoJul 27, 2026 • new comments - 🚨[wav2vec2] Support attn_implementation=sdpa dispatch
#46196 commented on
17 hours agoJul 29, 2026 • new comments - adding amd quark config class changes
#47322 commented on
38 minutes agoJul 30, 2026 • new comments - Add heterogeneous model support (per-layer config and modeling)
#45332 commented on
7 minutes agoJul 30, 2026 • new comments - Add new model `EfficientVit-SAM` to transformers
#47275 commented on
2 days agoJul 28, 2026 • new comments - Improve Trainer DataLoader Controls for Streaming and Multiprocessing
#47164 commented on
4 hours agoJul 29, 2026 • new comments - fix bugs for clvp model
#47127 commented on
2 days agoJul 27, 2026 • new comments - Enhance `get_balanced_memory` to ensure adequate GPU allocation for small models with large layers
#47426 commented on
2 hours agoJul 30, 2026 • new comments - Fix _is_packed_sequence for 1D and 3D position_ids
#47342 commented on
3 hours agoJul 29, 2026 • new comments - Batch Rebalance Data Sampler
#47340 commented on
yesterdayJul 28, 2026 • new comments - Route byte-level llama tokenizers to TokenizersBackend
#47017 commented on
4 days agoJul 26, 2026 • new comments - OpenVINO HF Exporter
#47003 commented on
2 days agoJul 27, 2026 • new comments - TP dtensor
#47194 commented on
2 days agoJul 28, 2026 • new comments - Also size the device_map buffer from the largest leaf module (follow-up to #47203)
#47211 commented on
yesterdayJul 29, 2026 • new comments - pytorch 2.13 is released , updating transformer to handle the fix tf32 issue
#47222 commented on
5 days agoJul 25, 2026 • new comments - Respect min_tokens_to_keep in TopHLogitsWarper
#47262 commented on
last weekJul 23, 2026 • new comments - Pipeline parallel naive inference
#47289 commented on
2 days agoJul 28, 2026 • new comments - [DeepSpeed] Composite models are not initialized correctly under ZeRO-3
#47297 commented on
last weekJul 23, 2026 • new comments - Experiments with Neuron
#47306 commented on
5 days agoJul 25, 2026 • new comments - Add chat input sanitization
#47386 commented on
19 hours agoJul 29, 2026 • new comments - fix(trainer): fix resume_from_checkpoint logic bug (issue #47375)
#47391 commented on
last weekJul 23, 2026 • new comments - Add LoMa keypoint matching model
#47401 commented on
yesterdayJul 28, 2026 • new comments - Make _added_tokens_decoder the single source of truth for added-token state
#47440 commented on
2 days agoJul 28, 2026 • new comments - Fix ESMFold error in fp16 precision
#47472 commented on
last weekJul 23, 2026 • new comments - Fix CLVP `c_proj` residual scaled-init silently never applied
#47478 commented on
10 hours agoJul 29, 2026 • new comments - Chat template inconsistencies in tool-calling support
#45419 commented on
last weekJul 23, 2026 • new comments - ESMFold throws error when explicitly cast to fp16
#47470 commented on
last weekJul 23, 2026 • new comments - generation error when same token in "forced_eos_token_id" and "supress_token" parameter.
#24099 commented on
last weekJul 24, 2026 • new comments - MetalConfig quantization: batched generate silently corrupts all rows after the first — affine_qmm_t computes only batch element 0 of 3D inputs (verified, one-line fix)
#47331 commented on
5 days agoJul 24, 2026 • new comments - Return updated attention mask from Wav2Vec 2.0
#25307 commented on
5 days agoJul 24, 2026 • new comments - Added-token encoder cache goes stale when tokenizers mutate _added_tokens_decoder directly, giving ids that do not round-trip
#47439 commented on
5 days agoJul 25, 2026 • new comments - YOLO model failing on multi GPU device map auto
#46869 commented on
4 days agoJul 25, 2026 • new comments - Unable to resume model training with a locally downloaded model and trust remote code
#46165 commented on
4 days agoJul 25, 2026 • new comments - Is transformers 5.12.1 requires tokenizers<=0.23.0,>=0.22.0 really right?
#46785 commented on
4 days agoJul 26, 2026 • new comments - [BUG] FP8 model cannot inference on B200 (SM>=100)
#46209 commented on
2 days agoJul 27, 2026 • new comments - tokenizers<=0.23.0 upper bound blocks the only existing 0.23.x release
#47429 commented on
2 days agoJul 28, 2026 • new comments - Batch size schedulers
#31222 commented on
2 days agoJul 28, 2026 • new comments - ZeRO-3 zero.Init does not partition composite minimax_m3_vl language submodule -> OOM on multi-GPU load
#46822 commented on
yesterdayJul 29, 2026 • new comments - Allow to give the dataset multiprocessing_context
#34793 commented on
yesterdayJul 29, 2026 • new comments - Please update `tokenizers` version check
#45736 commented on
11 hours agoJul 29, 2026 • new comments - Checkpoint conversion silently loads transposed fused-expert weights when the swapped dims are equal (Transpose(check_dims=True), qwen3_vl_moe)
#47031 commented on
9 hours agoJul 29, 2026 • new comments - FineGrainedFP8 DeepGEMM path returns silently wrong results on SM100 (UE8M0 scale rounding without requantization)
#47030 commented on
8 hours agoJul 29, 2026 • new comments - Remove device to host sync triggered in _flash_attention_forward
#39213 commented on
4 hours agoJul 29, 2026 • new comments - Mamba Models - Missing MambaForSequenceClassification
#30431 commented on
4 hours agoJul 29, 2026 • new comments - SentencePieceBackend should decode byte-fallback tokens
#47473 commented on
4 hours agoJul 29, 2026 • new comments - [Gemma3N] Not able to add new special tokens to model/tokenizer due to projection error
#39921 commented on
4 hours agoJul 29, 2026 • new comments - Add Molmo2
#43451 commented on
1 hour agoJul 30, 2026 • new comments - add HyperClovaX Vision
#44314 commented on
yesterdayJul 29, 2026 • new comments - 🚨🚧 FeatureExtractor → AudioProcessor
#44394 commented on
yesterdayJul 28, 2026 • new comments - Add Audio-Visual Flamingo model
#45586 commented on
2 days agoJul 28, 2026 • new comments - Enable MTP support for Qwen3.5
#45638 commented on
yesterdayJul 28, 2026 • new comments - [generation] Encode multimodal data only once
#45783 commented on
3 hours agoJul 29, 2026 • new comments - LocateAnything 3B
#46300 commented on
4 days agoJul 25, 2026 • new comments - 🚨 [DiffusionGemma] Return router logits and load balancing loss
#46642 commented on
2 days agoJul 27, 2026 • new comments - Fix CPM-Ant LM head tying (vocabulary-slice tie)
#46835 commented on
3 days agoJul 27, 2026 • new comments