Hi! As you know, like many other projects the Android Open-Source Project was affected by the recent kernel.org downtime. So, we’re pleased to let you know that the Gingerbread source code is now available again, and AOSP git servers are back online. Even before the kernel.org downtime, it was clear that AOSP was sometimes taxing kernel.org’s git infrastructure. When we did the Gingerbread source release, for example, load due to AOSP made part of kernel.org unusable for several days. This isn’t fair to kernel.org’s staff or the community, so for some time we’ve been preparing our own git hosting on Google servers. We were finishing up just as kernel.org experienced their downtime, so the Gingerbread source is now available on Google’s servers. Accordingly, the git URLs have changed. Here are the instructions to access the new git servers: You need to get the latest version of the repo tool: curl https://dl-ssl.google.com/dl/googlesource/git-repo/repo > ~/bin/repo You need to initialize a new repository: repo init -u https://android.googlesource.com/platform/manifest -b android-2.3.7_r1 The full instructions are at http://source.android.com/source/downloading.html There are a few limitations to be aware of: Our priority has been getting the main source code mirrors back online, so for the moment gitweb source browsing and Gerrit Code Review are still unavailable. We are now working on bringing AOSP’s Gerrit Code Review site back up, and hope to be able to say something here soon. It might be a little while longer before gitweb comes back, unfortunately, since Gerrit Code Review is the next priority. To reiterate, these servers contain only the ‘gingerbread’ and ‘master’ branches from the old AOSP servers. We plan to release the source for the recently-announced Ice Cream Sandwich soon, once it’s available on devices. As these new servers are, well, new, there may be hiccups if we encounter unexpected issues. However we’re keeping a close eye on them and will respond to any issues as quickly as possible. Finally, we’d like to send a huge “thank-you” to the kernel.org community and Oregon State University Open-Source Lab staff. They’ve done an incredible job hosting the AOSP source code mirror and Gerrit Code Review for nearly 3 years. Without them, it’s safe to say that AOSP would not be where we are today. Thanks, and happy coding! - Dan
Как мы запустили Qwen3.8–27B целиком на RTX 5060 8 GB и получили ~30 токенов/с
Наш проект называется ExVRAM Lab. Это открытая исследовательская лаборатория, в которой мы проверяем, насколько большие локальные LLM можно запускать на обычных видеокартах с ограниченным объёмом VRAM, если использовать ultra‑low‑bit quantization, полное размещение весов на GPU и существующие open‑source inference‑технологии.
ExVRAM расшифровывается как Exchange Compute for VRAM. Основная идея проекта — в ряде сценариев выгоднее потратить часть свободной вычислительной мощности GPU на работу с более компактным представлением весов, чем хранить часть модели в оперативной памяти и постоянно передавать данные через PCIe.
Когда мы начинали проект, исходный вопрос был достаточно простой: можно ли запустить dense‑модель примерно на 27 миллиардов параметров на видеокарте всего с 8 ГБ VRAM так, чтобы она не просто «запустилась», а работала полностью на GPU, поддерживала длинный контекст и обеспечивала нормальную интерактивную скорость генерации.
В качестве основной тестовой системы мы используем NVIDIA GeForce RTX 5060 8 GB на архитектуре Blackwell. Основная модель в текущих экспериментах — Qwen3.8–27B.
Сначала результат выглядел не слишком впечатляюще. Модель запускалась, но значительная часть весов оставалась в системной памяти. Около 5,7 GiB весов находилось на GPU, ещё примерно 2,5 GiB — на CPU. Скорость генерации составляла порядка 3,8–5,5 токена в секунду.
При этом сама видеокарта была загружена далеко не полностью.
Это стало одним из первых важных наблюдений проекта. Проблема заключалась не столько в нехватке вычислительной мощности RTX 5060, сколько в том, что часть decoder weights находилась в RAM. Во время autoregressive generation данные приходилось постоянно передавать между CPU и GPU через PCIe.

