Articles & Guides

Multimodal

Claude API relay guides, detection insights and hands-on LLM API benchmarks

34 articles

CCTest · Blog
Black Forest Labs unveils Flux3, a multimodal model built for native audio-video generation
Multimodal
cctest.ai
Multimodal

Black Forest Labs unveils Flux3, a multimodal model built for native audio-video generation

Germany-based AI startup Black Forest Labs has introduced Flux3, a multimodal foundation model designed to unify image, video, audio and action representations. The model is described as using a Self-Flow architecture and supporting synchronized audio-video generation up to 20 seconds.

Read more
CCTest · Blog
KnowAct-GUIClaw: A Self-Evolving GUI Assistant Built on Memory and Skills
Multimodal
cctest.ai
Multimodal

KnowAct-GUIClaw: A Self-Evolving GUI Assistant Built on Memory and Skills

KnowAct-GUIClaw introduces a “Know Deeply, Act Perfectly” paradigm for personal GUI assistants, aiming to address OpenClaw’s limitations in cross-platform GUI interaction and self-evolution. The framework combines experience-based memory, a self-evolving skill library, and reflection to improve task execution over time.

Read more
CCTest · Blog
FM²: A federated foundation model framework for heterogeneous multimodal medical imaging
Multimodal
cctest.ai
Multimodal

FM²: A federated foundation model framework for heterogeneous multimodal medical imaging

FM² tackles a practical barrier in medical AI: foundation models need data from many institutions, but privacy rules often prevent central aggregation. The paper focuses on modality heterogeneity across clients and proposes a unified federated training framework.

Read more
CCTest · Blog
Thinking Machines Releases Inkling, an Open Multimodal Model Near the Trillion-Parameter Scale
Multimodal
cctest.ai
Multimodal

Thinking Machines Releases Inkling, an Open Multimodal Model Near the Trillion-Parameter Scale

Thinking Machines has released Inkling on Hugging Face, an open multimodal model designed to natively accept image, text, and audio inputs. With a 1M-token context window and a sparse MoE architecture, it targets advanced multimodal reasoning and downstream adaptation.

Read more