| PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models | Jan 10, 2024 | GPUImage Generation | CodeCode Available | 7 | 5 |
| Large Language Model Agent: A Survey on Methodology, Applications and Challenges | Mar 27, 2025 | Language ModelingLanguage Modelling | CodeCode Available | 7 | 5 |
| Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers | May 9, 2024 | | CodeCode Available | 7 | 5 |
| SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild | Mar 24, 2025 | Instruction FollowingMath | CodeCode Available | 7 | 5 |
| Logo-LLM: Local and Global Modeling with Large Language Models for Time Series Forecasting | May 16, 2025 | Time SeriesTime Series Forecasting | CodeCode Available | 7 | 5 |
| DragAnything: Motion Control for Anything using Entity Representation | Mar 12, 2024 | ObjectVideo Generation | CodeCode Available | 7 | 5 |
| EvoRL: A GPU-accelerated Framework for Evolutionary Reinforcement Learning | Jan 25, 2025 | BenchmarkingEvolutionary Algorithms | CodeCode Available | 7 | 5 |
| Efficient MedSAMs: Segment Anything in Medical Images on Laptop | Dec 20, 2024 | Image SegmentationMedical Image Segmentation | CodeCode Available | 7 | 5 |
| Aligning Anime Video Generation with Human Feedback | Apr 14, 2025 | Video Generation | CodeCode Available | 7 | 5 |
| Chronos: Learning the Language of Time Series | Mar 12, 2024 | Gaussian ProcessesLanguage Modeling | CodeCode Available | 7 | 5 |
| Adding Conditional Control to Text-to-Image Diffusion Models | Feb 10, 2023 | Image GenerationLayout-to-Image Generation | CodeCode Available | 7 | 5 |
| OASIS: Open Agent Social Interaction Simulations with One Million Agents | Nov 18, 2024 | Large Language ModelRecommendation Systems | CodeCode Available | 7 | 5 |
| Muon is Scalable for LLM Training | Feb 24, 2025 | Computational Efficiency | CodeCode Available | 7 | 5 |
| An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents | May 21, 2025 | Reinforcement Learning (RL) | CodeCode Available | 7 | 5 |
| Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model | Jun 10, 2025 | Language ModelingLanguage Modelling | CodeCode Available | 7 | 5 |
| EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions | Jul 11, 2024 | Image Animation | CodeCode Available | 7 | 5 |
| CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases | Aug 7, 2024 | HumanEvalmbpp | CodeCode Available | 7 | 5 |
| Adaptive In-conversation Team Building for Language Model Agents | May 29, 2024 | DiversityLanguage Modeling | CodeCode Available | 7 | 5 |
| Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models | Mar 8, 2023 | | CodeCode Available | 7 | 5 |
| MiniMax-01: Scaling Foundation Models with Lightning Attention | Jan 14, 2025 | Mixture-of-Experts | CodeCode Available | 7 | 5 |
| Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold | May 18, 2023 | Image ManipulationPoint Tracking | CodeCode Available | 7 | 5 |
| MeshAnything: Artist-Created Mesh Generation with Autoregressive Transformers | Jun 14, 2024 | Decoder | CodeCode Available | 7 | 5 |
| BricksRL: A Platform for Democratizing Robotics and Reinforcement Learning Research and Education with LEGO | Jun 25, 2024 | reinforcement-learningReinforcement Learning | CodeCode Available | 7 | 5 |
| EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees | Jun 24, 2024 | | CodeCode Available | 7 | 5 |
| DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference | Jan 9, 2024 | BenchmarkingText Generation | CodeCode Available | 7 | 5 |
| MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis | Dec 19, 2024 | Audio GenerationAudio Synthesis | CodeCode Available | 7 | 5 |
| AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era | Dec 13, 2024 | Image to Video GenerationVideo Generation | CodeCode Available | 7 | 5 |
| Better than classical? The subtle art of benchmarking quantum machine learning models | Mar 11, 2024 | BenchmarkingBinary Classification | CodeCode Available | 7 | 5 |
| Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant | Oct 20, 2024 | Question Answeringspeech-recognition | CodeCode Available | 7 | 5 |
| GenAD: Generalized Predictive Model for Autonomous Driving | Mar 14, 2024 | Autonomous Drivingmodel | CodeCode Available | 7 | 5 |
| FAST-LIVO2: Fast, Direct LiDAR-Inertial-Visual Odometry | Aug 26, 2024 | NeRFState Estimation | CodeCode Available | 7 | 5 |
| aiXcoder-7B: A Lightweight and Effective Large Language Model for Code Processing | Oct 17, 2024 | AttributeCode Completion | CodeCode Available | 7 | 5 |
| AgentOrchestra: A Hierarchical Multi-Agent Framework for General-Purpose Task Solving | Jun 14, 2025 | | CodeCode Available | 7 | 5 |
| MAGI-1: Autoregressive Video Generation at Scale | May 19, 2025 | Video Generation | CodeCode Available | 7 | 5 |
| Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding | May 14, 2024 | Image GenerationLanguage Modeling | CodeCode Available | 7 | 5 |
| ComfyUI-Copilot: An Intelligent Assistant for Automated Workflow Development | Jun 5, 2025 | Large Language Model | CodeCode Available | 7 | 5 |
| Kimi-Audio Technical Report | Apr 25, 2025 | Audio Question AnsweringQuestion Answering | CodeCode Available | 7 | 5 |
| Bilateral Reference for High-Resolution Dichotomous Image Segmentation | Jan 7, 2024 | Camouflaged Object SegmentationDichotomous Image Segmentation | CodeCode Available | 7 | 5 |
| EvoGP: A GPU-accelerated Framework for Tree-based Genetic Programming | Jan 21, 2025 | Feature EngineeringGPU | CodeCode Available | 7 | 5 |
| AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems | Mar 9, 2025 | | CodeCode Available | 7 | 5 |
| StarCoder 2 and The Stack v2: The Next Generation | Feb 29, 2024 | Code CompletionCode Generation | CodeCode Available | 7 | 5 |
| Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming | Aug 29, 2024 | Speech Synthesis | CodeCode Available | 7 | 5 |
| Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems | Dec 12, 2024 | | CodeCode Available | 7 | 5 |
| Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers | Jan 5, 2023 | In-Context LearningLanguage Modeling | CodeCode Available | 7 | 5 |
| DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | Oct 16, 2024 | | CodeCode Available | 7 | 5 |
| Intent-based Prompt Calibration: Enhancing prompt optimization with synthetic boundary cases | Feb 5, 2024 | Prompt Engineering | CodeCode Available | 7 | 5 |
| Improving Sample Quality of Diffusion Models Using Self-Attention Guidance | Oct 3, 2022 | DenoisingDiversity | CodeCode Available | 7 | 5 |
| EasyAnimate: A High-Performance Long Video Generation Method based on Transformer Architecture | May 29, 2024 | Image GenerationVideo Generation | CodeCode Available | 7 | 5 |
| HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters | May 26, 2025 | Human Animation | CodeCode Available | 7 | 5 |
| MagicQuill: An Intelligent Interactive Image Editing System | Nov 14, 2024 | Language ModelingLanguage Modelling | CodeCode Available | 7 | 5 |