Open source implementation of a vision transformer that can understand Videos using max vit as a foundation.
Repositories
kyegomez repositories
Open source scripts to create large scale datasets with rich detail for multi-modal models
Implementation of VisionLLaMA from the paper: "VisionLLaMA: A Unified LLaMA Interface for Vision Tasks" in PyTorch and Zeta
Implementation of Vision Mamba from the paper: "Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model" It's 2.8x faster than DeiT and saves 86.8% GPU memory when performing batch inference to extract features on high-res images
An plug in and play pipeline that utilizes segment anything to segment datasets with rich detail for downstream fine-tuning on vision models like CLIP, ViT, Imagebind, and so on!
Open source implementation of "Vision Transformers Need Registers"
Transformers + Mambas + LSTMS All in One Model
WARP: Warp Speed Protocol Commit Message ShortHand Framework
The world's first fully automated VC fund.
Implementation of the Paper: "Zamba: A Compact 7B SSM Hybrid Model" in Pytorch
Python SDK for agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks like CrewAI, Langchain, and Autogen
An Definitive and Unified AI Agents Framework to Automate Anything and Everything
Agora landing page
Building and customizing your own version of AI town - a virtual town where AI characters live, chat and socialize.
This collection brings together the highest-signal research papers in modern AI from the invention of the Transformer to the frontier work of 2024–2025 into a single, curated map of the field
Scripts for harvesting all the data from libgen for massive AI pretraining
A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.
Implementation of Alphafold 3 in Pytorch
APAC Ventures website
A clean, single-file PyTorch implementation of Attention Residuals (Kimi Team, MoonshotAI, 2026), integrated with Grouped Query Attention (GQA), SwiGLU feed-forward networks, and Rotary Position Embeddings (RoPE).