Scott Boland is challenging the established 'big three' of Australian fast bowling as selectors face a daunting schedule of 21 upcoming Test matches. With all primary quicks now fit, Australia's leadership must balance individual form, injury history, and the necessity of squad rotation.
Leads Brand Connect has officially announced the appointment of Raksha Narayan Kotian as the new Business Head for Coração Do Vale. This strategic leadership hire aims to drive growth and strengthen the brand's market positioning within the competitive lifestyle and hospitality sector.
FlowMimic introduces a novel framework for mask-free video editing by leveraging pixel-pair temporal warped flow fields to generate training data from image-based samples. By aligning image and video modalities through mutual imitation, the system internalizes editing capabilities, removing the need for external masks or auxiliary models.
AMC Entertainment has reported its second-quarter financial results, successfully exceeding analyst expectations for both earnings and revenue. This performance highlights the company's ongoing efforts to navigate a challenging theatrical landscape and manage its debt profile.
Australia has announced a 13-man squad for the upcoming two-Test series against Bangladesh, featuring the return of captain Pat Cummins, Josh Hazlewood, and Nathan Lyon. The selection marks a significant boost for the team as they prepare for a demanding 11-month schedule of international cricket.
A new report from CompTIA reveals that the technology sector is significantly contributing to national job growth, effectively bucking broader economic trends. This surge in demand for tech talent underscores the critical role of digital transformation across diverse industries.
JoyNexus introduces a service-oriented architecture for Vision-Language-Action (VLA) model post-training, moving away from exclusive resource allocation. By decoupling training, inference, and environment services, it enables efficient multi-tenancy and resource sharing for complex robotic workloads.
India and Spain have committed to accelerating technical discussions to link India's Unified Payments Interface (UPI) with Spain’s Bizum platform. This initiative, discussed during Commerce Minister Piyush Goyal’s recent European tour, aims to streamline cross-border transactions and deepen economic cooperation between the two nations.
President Droupadi Murmu has officially invited Moldovan businesses to explore investment opportunities in India during a high-level business forum. The visit underscores a strategic push to deepen bilateral economic cooperation and foster stronger trade ties between the two nations.
Mumbai-based Sagar Pictures Entertainment has announced a strategic shift from traditional film production to a global intellectual property powerhouse. The company plans to leverage its seven-decade legacy to develop multi-platform franchises, starting with a seven-part cinematic adaptation of the Shrimad Bhagavatam.
New research challenges the industry's reliance on peak success rates for AI security agents, proposing a cost-aware evaluation framework. The findings highlight that offensive and defensive agents exhibit fundamentally different scaling behaviors, requiring developers to prioritize operational efficiency over raw reasoning budgets.
Researchers have introduced VideoRAE, a novel representation autoencoder that leverages frozen Video Foundation Models to enhance generative video modeling. By compressing hierarchical features, the system achieves superior reconstruction and significantly faster training speeds compared to traditional 3D-VAE architectures.
A recent industry report indicates that recruitment for artificial intelligence roles in India is significantly outpacing the growth of the broader IT sector. This shift underscores a strategic pivot among Indian technology firms as they prioritize generative AI and machine learning capabilities to maintain global competitiveness.
The Central Board of Secondary Education and the National Council of Educational Research and Training have initiated orientation workshops to prepare educators for the rollout of new Class 9 mathematics and science textbooks. These sessions are designed to align classroom instruction with the pedagogical goals of the National Education Policy 2020.
Vision-language-action models often struggle with the discrepancy between camera-frame visual input and robot-frame action output. The introduction of robot-centric pointmaps offers a solution by encoding 3D scene data directly in a robot-relative coordinate system, enhancing generalization across diverse camera setups.
Union Commerce Minister Piyush Goyal concluded a strategic four-nation tour of Europe, focusing on accelerating the implementation of the India-EU Free Trade Agreement. The visit underscored India's ambition to become a global manufacturing and technology hub through enhanced bilateral cooperation in digital finance, clean energy, and advanced manufacturing.
A new analysis from CompTIA indicates that technology job postings have reached their highest level in three years, driven largely by the surging demand for artificial intelligence expertise. This shift marks a significant turnaround for the sector following a period of post-pandemic contraction.
Researchers have introduced Audio-Visual Flamingo (AV-Flamingo), an open-source large language model designed to master complex, long-form audio-visual reasoning. By utilizing a massive new dataset and a specialized three-stage training curriculum, the model sets a new standard for temporal alignment and interpretability in multimodal AI.
A groundbreaking study has identified the specific type of meteorite responsible for the mass extinction 66 million years ago. By analyzing isotopic signatures, researchers have confirmed the impactor was a rare carbonaceous chondrite originating from the outer solar system.
The Afghanistan Chamber of Commerce and Investment has initiated high-level talks with Indian officials to bolster bilateral trade and investment. The move comes as Kabul seeks to mitigate the impact of declining transit trade with Pakistan by leveraging alternative routes like the Chabahar port.
Recent research explores the application of the Muon optimizer in sparse-reward agentic reinforcement learning, demonstrating significant performance gains over traditional AdamW. By optimizing hidden weight matrices within specific policy frameworks, Muon accelerates convergence and improves success rates in complex task environments like ALFWorld.
India’s IT sector has recorded its lowest level of active job openings in 28 months, with current vacancies standing at 93,000. This sharp decline reflects a broader trend of caution among major tech employers as they navigate global economic uncertainty and shifting client demands.
Researchers have utilized chess as a controlled testbed to analyze how pretraining choices influence the effectiveness of reinforcement learning in large language models. The study reveals that pretraining loss is a strong predictor of post-RL performance, offering new insights into the science of model reasoning.
India's technology sector has experienced a cooling in hiring activity, with active job openings dropping by 8% in April 2026. This decline marks the second-lowest start to a fiscal year in six years, signaling a shift toward cautious recruitment strategies.
Moonshot AI is transitioning its Kimi CLI, a terminal-based AI agent for software development, to the new Kimi Code CLI. This evolution introduces enhanced features for code editing, shell command execution, and broader IDE integration, while maintaining existing functionalities.
Unsloth has released Inkling-GGUF, a general-purpose multimodal model capable of processing text, image, and audio inputs to generate text outputs. This model, featuring a sparse Mixture-of-Experts architecture, is designed for developers building AI applications such as agentic systems, coding assistants, and chatbots. It supports local deployment via several open-source libraries and offers multilingual capabilities.
Robbyant Team has introduced LingBot-Map, a feed-forward 3D foundation model designed for real-time streaming 3D reconstruction. It leverages a Geometric Context Transformer to unify coordinate grounding, dense geometric cues, and long-range drift correction within a single framework. The model demonstrates high-efficiency streaming inference and state-of-the-art reconstruction performance on various benchmarks.
Code-review-graph is a new tool designed to significantly reduce token consumption and improve the accuracy of AI-powered code reviews. By building a local, persistent structural map of a codebase, it provides AI assistants with precise context, focusing reviews only on relevant changes and their blast radius. This approach aims to make AI coding tools more efficient and cost-effective, particularly in large repositories.
A new paper critically examines the evaluation protocols for automatic harness evolution in LLM agents. It highlights concerns regarding potential overfitting to benchmarks and the need for fairer comparisons against simpler test-time scaling methods under matched computational budgets. The research suggests that current harness evolution methods may not consistently outperform these baselines and exhibit limited generalization.
Google's Protocol Buffers (protobuf), a language-neutral, platform-neutral, and extensible mechanism for serializing structured data, has seen recent updates including enhanced Bazel support with Bzlmod and updated GCC testing. These changes aim to improve build stability, compatibility, and development workflows for protobuf users and contributors. The project continues to emphasize working from supported releases for optimal stability.
DocuSeal is an open-source platform designed to provide secure and efficient digital document signing and processing, serving as an alternative to proprietary solutions like DocuSign. It enables users to create, fill, and sign PDF forms online with a mobile-optimized web tool. The platform offers a range of features for developers and businesses, including API integrations and various deployment options.
TurboQuant, a new Rust-based vector index with Python bindings, leverages Google Research's TurboQuant algorithm to significantly reduce memory footprint and improve search speeds compared to FAISS. It achieves up to 16x compression, enabling a 10 million document corpus to fit into 4 GB of RAM, while offering faster search times on ARM and competitive performance on x86 architectures.
New research highlights that current global vision models struggle with out-of-distribution generalization, similar to language models. The study demonstrates that a combination of local, foveated perception and recurrent neural networks is crucial for robust compositional generalization in visual reasoning tasks. This approach offers significant accuracy improvements over brute-force scaling of global models.
GnLOLot has released a new GGUF quantized model, MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF, designed for efficient local deployment and text generation tasks. This model integrates with a wide array of local AI tools and libraries, including `llama.cpp`, `llama-cpp-python`, vLLM, Ollama, and Unsloth Studio, facilitating accessible development for AI engineers.
Researchers have introduced RoboTTT, a novel robot model and training methodology that extends visuomotor context to 8,000 timesteps, a three-order-of-magnitude increase over prior state-of-the-art. This advancement enables new capabilities such as one-shot in-context imitation and on-the-fly policy improvement without increasing inference latency. RoboTTT integrates Test-Time Training into robot foundation models, demonstrating significant performance gains on complex real-robot manipulation tasks.
SearchOS introduces a novel system-level multi-agent framework designed to enhance the robustness of open-domain information-seeking agents. It addresses the common problem of agents getting trapped in repetitive search loops by externalizing search progress into explicit, persistent, and shared state. This framework leverages a Search-Oriented Context Management (SOCM) system and a pipeline-parallel scheduling mechanism to improve efficiency and accuracy in information retrieval.
Researchers have introduced Hierarchical Denoising for Visual Reasoning (HDR), a novel framework designed to enhance multi-step reasoning in video models. HDR integrates hierarchical latents into causal video generation, enabling coarse-to-fine reasoning and addressing limitations in logical consistency and low-latency streaming found in existing diffusion models. This approach significantly improves success rates and reasoning consistency across complex visual tasks.
New research reveals that Large Language Models (LLMs) frequently violate basic statistical self-consistency principles, particularly when aggregating information from partitioned data. The study introduces the 'macro fallacy,' where fine-grained subpopulation estimates often yield more accurate aggregate results than direct population-level estimates. These findings highlight a critical gap in LLM's ability to reliably propagate subpopulation knowledge into broader statistical summaries.
MeanFlowNFT introduces a novel framework that integrates forward-process reinforcement learning (RL) with MeanFlow generators, which are known for their efficient few-step sampling. By bridging the gap between instantaneous velocity optimization in DiffusionNFT and average velocity sampling in MeanFlow, MeanFlowNFT enables reward-based alignment of these fast generators. This innovation leads to improved performance in image and video generation, often surpassing multi-step RL-tuned diffusion models with significantly fewer sampling steps.
A new paper, 'From Pixels to States: Rethinking Interactive World Models as Game Engines,' examines the potential of video generative models to power next-generation interactive game worlds. It analyzes the challenges in achieving true interactivity, persistence, and real-time generation, proposing a framework based on the traditional action-state-observation loop. The authors also introduce a scalable data engine for Black Myth: Wukong to support state-aware game world modeling.
Researchers have introduced BadWAM, a new framework for evaluating World-Action Drift Attacks against World-Action Models (WAMs). These attacks use subtle visual perturbations to desynchronize a WAM's imagined future from its executed actions, challenging the assumption of inherent robustness in these embodied AI systems. BadWAM highlights critical vulnerabilities, demonstrating how WAMs can 'dream right but act wrong' under adversarial conditions.
Wan-Dancer-14B, a new model from Wan-AI, introduces a hierarchical framework for generating long-duration, high-quality, and rhythmically coherent dance videos from music. This method decouples the generation process into global keyframe planning and local temporal refinement, ensuring structural and temporal continuity over minute-scale videos. The model and inference code are now available on Hugging Face, enabling developers to create diverse dance styles from input music and a reference image.
Prism ML has released Bonsai 27B, a 27B-class language model featuring binary transformer weights, enabling it to run on high-end smartphones like the iPhone 17 Pro Max. This model achieves a 14.2x reduction in size compared to FP16 while retaining approximately 90% of its intelligence, making advanced LLM capabilities accessible on edge devices.
Large language model (LLM) agents are increasingly automating machine learning engineering (MLE), but medical imaging presents unique challenges due to its modality-specific experimentation and stringent validation requirements. A new framework, AMID (Autonomous Multi-Agent framework for medical Imaging model Development), addresses these by introducing Data-Conditioned Method Planning and Verification-Guided Two-Stage Optimization. This system aims to transform bespoke manual medical imaging model development into an agentic workflow, producing high-performing and auditable model artifacts.
Researchers have introduced SynthDocBench, a novel synthetic benchmark designed to systematically evaluate Vision Language Models (VLMs) on long-context visual document understanding. This benchmark controls for factors like document length, layout complexity, and modality, uncovering specific failure modes in frontier VLMs that existing benchmarks do not address. The findings suggest that current models may be overfitting to benchmark artifacts rather than achieving robust long-context understanding.
A new framework called RGB In and RGB Out (RINO) proposes a unified approach for diverse vision tasks by representing all visual information as RGB images and converting tasks into RGB-to-RGB image editing problems. This paradigm allows a single model to handle various visual tasks through a shared visual interface, analogous to how large language models process text. RINO demonstrates robust zero-shot performance across dense understanding and dense-conditioned generation tasks without task-specific fine-tuning.
Current visual generators struggle with world-knowledge, often fabricating details for requests outside their training data. New research introduces a 'teach-then-search' co-training framework to dynamically identify and evolve a generator's knowledge boundary, enabling more accurate and grounded visual outputs for long-tail, evolving user prompts. This approach aims to improve agentic visual generation by intelligently integrating external search tools.
Researchers have introduced ChartCynics, an agentic dual-path framework designed to improve Vision-Language Models' (VLMs) ability to interpret misleading charts. By decoupling perception from verification and employing a skeptical reasoning paradigm, ChartCynics addresses deceptive visual structures and distorted data representations. This approach has demonstrated significant performance improvements over existing VLM backbones, establishing a new foundation for trustworthy chart interpretation.
Thinking Machines has released Inkling, a general-purpose multimodal model capable of processing text, image, and audio inputs to generate text outputs. Designed for developers, Inkling features an open-weights architecture with a sparse Mixture-of-Experts (MoE) backbone, supporting a range of AI applications from agentic systems to chatbots. The model is available for local deployment via several open-source libraries and offers multilingual capabilities.
Existing benchmarks for scientific data analysis often overlook the diverse types of scientific claims LLMs need to support. SDABench reorients evaluation around six core capabilities across five scientific domains, providing a more granular assessment of LLM performance in scientific discovery. Initial evaluations reveal that while LLMs handle descriptive analysis well, they struggle with tasks requiring complex reasoning such as assumption selection and mechanistic modeling.
Researchers have developed HealthClaw, an open-source AI agent architecture designed for longitudinal personal health management. Unlike traditional health AI systems that process requests in isolation, HealthClaw features a self-evolving memory that adapts to a person's changing routines, preferences, and health data over time. This architecture significantly improves answer accuracy and privacy while reducing context exposure in health support scenarios.
Researchers have introduced Deep Interaction, a new method designed to efficiently correct reasoning errors in large language models (LLMs) by allowing direct editing of erroneous Chain-of-Thought (CoT) steps. This approach refines the corrected CoT into a distilled prompt, guiding the LLM along an accurate reasoning path. Experimental results demonstrate significant improvements in correction success rates and reduced token usage compared to existing methods.
Traditional penetration testing focuses on infrastructure compromise. A new paper proposes an expanded framework for AI-enabled systems, shifting focus to objective-driven behavioral evaluation. This redefinition accounts for adversarial influence on AI behavior without direct system compromise.
PalmClaw is an open-source agent framework designed to run natively on mobile phones, enabling LLM agents to directly access device capabilities. This approach bypasses the limitations of GUI-based mobile agents, offering improved task success and significantly reduced completion times. The framework manages sessions, memory, skills, tools, and the agent loop directly on the device, exposing device capabilities as structured tools.
Unsloth has released Qwen3.6-27B-NVFP4, an NVFP4 quantized version of the Qwen3.6-27B model, offering significantly faster throughput and improved agentic coding capabilities. This release focuses on stability and real-world utility, providing developers with a more responsive and productive coding experience, particularly for frontend workflows and repository-level reasoning. The model is compatible with various inference frameworks and supports Multi-Token Prediction for optimized decoding.
Researchers have introduced SIS-Bench, a new benchmark designed to evaluate multimodal large language models (MLLMs) in autonomous UAV systems. This benchmark addresses the critical gap in assessing an agent's self-awareness alongside its spatial cognition, crucial for complex real-world operations. Initial evaluations using SIS-Bench reveal current MLLMs exhibit limitations in dynamic, agent-centered processes, highlighting an imbalance between spatial understanding and self-awareness.
Researchers have introduced "generative compilation," a novel method that provides on-the-fly compiler feedback to AI models during code generation. This approach, centered around a "sealor" transformation, converts partial programs into diagnosable complete ones, enabling early error detection and improving the functional correctness of AI-generated code, particularly for languages like Rust.
New research reveals that applying length penalties in reinforcement learning for large language models (LLMs) can shorten chain-of-thought reasoning, but at the cost of reduced monitorability. While models maintain accuracy with fewer reasoning tokens, the influence of misleading hints becomes harder to detect. This creates a trade-off between computational efficiency and the transparency of an LLM's decision-making process.
SPEAR is a new Python library designed to enhance the generality, programmability, and rendering speed of photorealistic simulators for embodied AI research. It achieves this by providing programmatic control over any Unreal Engine application, exposing over 14,000 unique UE functions to Python and significantly improving rendering performance.
A new research paper introduces E3 (Estimate, Execute, Expand), a strategy designed to combat the common issue of LLM agents over-reading and re-processing information, which leads to significant inefficiencies. E3 enables agents to estimate task complexity, execute a minimum viable path, and expand scope only when necessary, resulting in substantial reductions in operational costs and resource consumption.
Researchers have introduced two dynamic resource allocation mechanisms for Ensemble Determinization Monte Carlo Tree Search (ED-MCTS), significantly improving its performance in high-uncertainty adversarial games. These enhancements, Dynamic Number of Determinizations and Dynamic Simulation Allocation, adapt search resources based on real-time search behavior and potential knowledge gain. The advancements were validated across popular tabletop games like Jaipur, Lost Cities, and Splendor.
A new methodology, Popperian Placebo-controlled Evaluation (PoPE), assesses whether frozen small code LLMs can operationally use error evidence for self-repair. The study found that error content, when compared against channel-specific placebos, did not demonstrate superior performance in either prompt-based or weight-adapter-based repair mechanisms. These results suggest that the specific content of error feedback may not be as effective as previously assumed for these models.
Remote sensing change detection (RSCD) models often struggle with catastrophic forgetting when adapting to new data domains. A new framework, DG-FDD, addresses this by integrating a Difference-Guided Dynamic Adapter and Frequency-Decoupled Knowledge Distillation to preserve bitemporal discrepancy cues and enable stable knowledge transfer without historical data. This approach significantly reduces performance degradation in incremental learning scenarios.
A recent paper conducts a principled analysis of deep reinforcement learning (DRL) evaluation and design, revealing that canonical paradigms have led to incorrect conclusions. The research demonstrates that the asymptotic performance of DRL algorithms does not exhibit a monotonic relationship between performance rankings and data-regimes, necessitating a re-evaluation of current research practices.
ATH-MaaS has introduced OvisOCR2, a new 0.8B parameter end-to-end model designed for page-level document parsing. Built on Qwen3.5-0.8B, OvisOCR2 excels at converting document images into structured Markdown, including text, formulas, tables, and visual regions, setting new benchmarks in the field. Its compact size and robust performance make it a significant development for multimodal AI applications.
Researchers have introduced SpectraReward, a training-free reward function that leverages pretrained Multimodal Large Language Models (MLLMs) as zero-shot reward models for text-to-image generation reinforcement learning. This method measures prompt recoverability from generated images using image-conditioned log-likelihood, eliminating the need for preference labels or reward model fine-tuning. A specialized version, Self-SpectraReward, enables unified multimodal models to self-improve without external reward models.
GnLOLot has released MiniCPM5-1B-Claude-Opus-Fable5-Thinking, a 1-billion parameter language model fine-tuned for improved coding and instruction-following capabilities. Built upon the MiniCPM5-1B base, this model integrates 'Thinking' chain-of-thought reasoning and supports a 128K context length, making it suitable for local and edge deployments. A V2.0 with enhanced tool-calling has also been released.
Current Video Large Language Models (Video LLMs) often act as black boxes, providing answers without verifiable visual evidence. Researchers have introduced Evidence-Backed Video Question Answering (E-VQA), a new task that requires models to output both a semantic answer and precise spatio-temporal evidence. This approach aims to enhance explainability and improve the visual perception capabilities of Video LLMs.
Xiaomi-Robotics-U0 is a 38-billion-parameter multimodal autoregressive model designed for unified embodied synthesis. It extends foundation image and video generation to embodied scenarios, addressing challenges like multi-view consistency and robot embodiment constraints. This model integrates various generative tasks, including text-to-image, image editing, and embodied video generation, while preserving the generalization capabilities of pre-trained world foundation models.
Researchers have introduced Latent-Identity Tuning, a novel method for precise facial editing within text-to-image personalization models. This approach modifies the latent representation of an identity directly, enabling consistent and diverse edits across generated images without requiring additional model training. By leveraging the latent space of a frozen encoder, the method identifies semantic directions for localized and fine-grained facial modifications.
Researchers have introduced MET (Multilingual Ethics with Theory-grounded reasoning), a novel approach to enhance language models' moral decision-making across diverse linguistic and cultural contexts. This method, along with a new benchmark MCLASH and a self-distillation technique MET-D, addresses critical limitations in existing multilingual moral reasoning systems.
A new paper offers the first comprehensive overview of metacognition in Large Language Models (LLMs), analyzing its foundations, current progress, and future opportunities. It taxonomizes the emerging field, summarizes technical advancements, and discusses methods for measurement, evaluation, elicitation, and application of metacognitive abilities in LLMs. The review aims to stimulate further research into making AI systems more capable, transparent, and reliable.
Empero AI has released Qwythos-9B-v2, an updated version of their 9B parameter language model built on the Qwen3.5 stack. This iteration primarily focuses on eliminating repetitive looping behavior and restoring the native multi-token-prediction (MTP) head, while preserving its 1M-token context and strong reasoning capabilities. The model remains intentionally uncensored for research and specialized technical applications.
Researchers have introduced AdvancedMathBench, a new benchmark suite designed to evaluate the advanced mathematical reasoning capabilities of large language models (LLMs). This suite addresses limitations in existing benchmarks by offering broader disciplinary coverage and more granular evaluation of proof generation and verification, extending to undergraduate and doctoral-level mathematics. Initial experiments reveal that even frontier models like GPT-5.5-xhigh still face significant challenges in these advanced mathematical tasks.
Robbyant has released LingBot-Video, the first open-source large-scale Mixture-of-Experts (MoE) video generation model specifically designed for embodied intelligence. This model aims to bridge the gap between video synthesis and real-world physical understanding, featuring an efficient MoE architecture and training on extensive embodied data.
Researchers have introduced MedPMC, an automated and continuously updatable framework designed to curate high-fidelity medical image-text pairs from permissively licensed literature. This framework addresses the critical shortage of high-quality, large-scale clinical data for training multimodal foundation models in medicine. MedPMC has demonstrated significant improvements in data quality and model performance across various medical tasks and benchmarks.
Flow-ERD is a novel multi-agent traffic simulator designed to enhance both the realism and diversity of simulated traffic scenarios, crucial for autonomous driving development. It employs a two-stage approach: Agent-Type Aware Flow Matching (AFM) for diverse, type-consistent motion generation, followed by Entropy-Regularized Distillation (ERD) to prevent mode collapse and mitigate covariate shift. This method addresses the current imbalance in traffic simulation benchmarks, which often prioritize realism over diversity.
A new research paper introduces a 'proactive memory agent' designed to address "behavioral state decay" in long-horizon AI tasks. This agent operates alongside an action agent, actively managing a structured memory bank and injecting relevant reminders to prevent critical information from being lost or overlooked. The plug-and-play module demonstrates improved performance across various benchmarks for both weaker and stronger action agents.
OpenCoF is a new framework designed to improve reasoning capabilities in video generation models through a novel Chain-of-Frame (CoF) approach. It features the OpenCoF-17K dataset and the Wan-CoF model, which leverage diverse temporal supervision and explicit reasoning tokens to enhance spatial and temporal understanding in generated videos. This framework aims to address the limitations of existing video generators that lack dedicated designs for complex reasoning tasks.
NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4, a deployment-optimized large language model derived from Nemotron-3-Super-120B-A12B. This model utilizes a hybrid Mixture-of-Experts (MoE) architecture with interleaved Mamba, MoE, and Attention layers, significantly improving inference efficiency for interactive and long-context workloads. It achieves this through Iterative Puzzle, a post-training compression framework, while maintaining strong downstream accuracy.
Recent research reveals that three prominent language model training methods—GRPO, Dr. GRPO, and DAPO—are fundamentally variations of a single mechanism. They all adjust a single metric: the standard deviation of sampled answers to a given prompt. This standard deviation directly correlates with the magnitude of the training update, indicating that disagreement among responses is a crucial driver of learning.
Researchers have developed PanoWorld, a new framework for panoramic world models that addresses long-range memory challenges by leveraging rotation-equivariant omnidirectional representations. PanoWorld simplifies camera trajectories and introduces a new dataset, World360, for evaluating physical consistency in diverse environments.
LongE2V, a novel approach, utilizes pre-trained video diffusion priors to address the challenges of event-based video reconstruction, prediction, and frame interpolation. By fine-tuning foundational video models, it achieves high data efficiency and superior perceptual quality, outperforming existing methods in temporal coherence and zero-shot generalization. The method introduces several key techniques to mitigate temporal drift and ensure precise consistency in long video sequences.
A new paper challenges the conventional text-only pretraining paradigm for large foundation models, demonstrating that directly leveraging visual documents without text extraction leads to superior performance. This 'Visual Pretraining' method consistently outperforms text-only pretraining across various backbones and benchmarks, offering a more efficient pathway to scalable language intelligence by incorporating rich visual cues often lost in text conversion.
ARDY is a novel streaming generation framework designed for high-fidelity, real-time 3D human motion synthesis. It addresses the limitations of existing methods by enabling interactive control via online text prompts and flexible kinematic constraints, crucial for animation, simulation, and robotics applications. ARDY achieves this through a hybrid representation and a two-stage autoregressive transformer denoiser.
Researchers have introduced Canvas360, a two-stage framework designed to enhance in-context panoramic generation. This framework leverages geometry-aware pretraining and task-specific fine-tuning, supported by a new large-scale dataset and novel modeling techniques. Canvas360 aims to improve geometric consistency and global coherence in generated panoramic images across various tasks.
Researchers have introduced UniClawBench, a novel benchmark designed to evaluate proactive AI agents in dynamic, real-world environments. Unlike previous benchmarks, UniClawBench focuses on five foundational model capabilities and uses live Docker containers for evaluation, providing a more robust assessment of agent performance.
Researchers have introduced SAM-MT, a novel framework built upon Segment Anything 2 (SAM2) that enables real-time interactive multi-target video segmentation. This approach addresses the limitations of traditional methods by decoupling latency from the number of targets, maintaining high frame rates even with multiple objects. SAM-MT achieves this through explicit target queries, decoupled masked attention, and sparse memory for temporal stability.
OPSD-V introduces an on-policy self-distillation paradigm to improve few-step autoregressive (AR) video diffusion models. By leveraging real long-video data for temporal context during training, OPSD-V mitigates error accumulation and weakened motion dynamics in long AR rollouts. This method enhances visual quality and motion dynamics without altering the original few-step inference path.
A new benchmark, IdeaGene-Bench (IG-Bench), has been introduced to evaluate AI systems' ability to understand and generate scientific ideas based on their evolutionary lineage. This framework models scientific concepts as 'Idea Genomes' that undergo inheritance, mutation, and recombination, similar to biological genomes. Initial experiments reveal a significant compositional bottleneck in current LLM-based systems, with the strongest performing at only 27.3% exact accuracy.
Unity MCP (Model Context Protocol) is an open-source project designed to integrate AI assistants directly with the Unity Editor. It provides LLMs with tools to manage assets, control scenes, edit scripts, and automate various game development workflows. This enables developers to leverage natural language interfaces for complex Unity tasks.
Superpowers is a new agentic skills framework and software development methodology designed for coding agents. It provides a structured approach to software development, emphasizing TDD, YAGNI, and DRY principles through a series of composable skills. The framework integrates with various coding agents like Claude Code, Antigravity, and GitHub Copilot CLI, guiding them from design specification to subagent-driven implementation and code review.
Researchers have introduced WorldSample, a novel framework designed to enhance real-robot reinforcement learning by integrating physical rollouts with high-fidelity synthetic transitions. This approach utilizes a real-synthetic loop, a post-trained world model, and Policy-Paced Learning to significantly reduce interaction costs and improve policy success rates in robot manipulation tasks. WorldSample addresses the limitations of traditional RL deployments on physical robots by generating realistic synthetic data and intelligently regulating its use.
GeoMix is a new descriptor-free 2D-3D matching framework that significantly improves visual localization accuracy by strengthening geometric discriminability. It introduces directional and distance-aware embeddings, learnable global context nodes, and a novel Mix-Training approach for multiple keypoint detectors. This advancement narrows the performance gap between descriptor-free and descriptor-based methods, offering benefits in privacy and map maintenance.
Herdr is a new terminal multiplexer designed specifically for managing multiple AI coding agents. It provides a real terminal environment for each agent, offers at-a-glance status updates (blocked, working, done, idle), and supports persistent sessions accessible from any terminal via SSH. Built in Rust, Herdr aims to streamline the developer workflow when orchestrating numerous AI assistants.
Researchers have introduced Exformer, an Extreme-Adaptive Transformer designed to improve time series forecasting, particularly for data containing rare but critical extreme events. This new framework addresses the limitations of traditional Transformer models that often underrepresent extreme patterns by treating all time points uniformly. Exformer incorporates a novel extreme-adaptive attention mechanism to explicitly model dependencies between normal and extreme events.
Yuxinlu1 has released Gemma4-12B v2, an updated GGUF model focused on agentic coding and tool-use capabilities. This iteration significantly improves performance on technical-agentic tasks compared to its base model, making advanced AI agent functionality accessible on local hardware with minimal VRAM requirements. The model is designed for multi-step technical tasks, debugging, and code generation.
Empero AI has released Qwythos-9B-Claude-Mythos-5-1M-GGUF, a quantized version of their 9B parameter reasoning model. This model features a 1M token context window, native function calling, and multimodal image input, making it suitable for local deployment on various GGUF-compatible runtimes.
New research challenges the prevailing understanding of performance improvements in self-alignment methods for diffusion transformers. Contrary to previous assumptions, the gains from methods like Self-Flow over SRA appear to stem primarily from data augmentation along the noise dimension, rather than interactions between tokens at different noise levels. The introduction of 'Attention Separation' demonstrates that blocking such interactions can even improve performance, highlighting the role of augmentation.