What if the model already knows how to solve many problems, but simply fails to pick the correct reasoning path consistently?
Applied AI. Gamer.
What if the model already knows how to solve many problems, but simply fails to pick the correct reasoning path consistently?
June 29 2026
The 100-year-old statistics model that quietly decides which ChatGPT answer you see.
June 22 2026
Can the model already generate correct solutions but simply fail to select them consistently? Or is the reasoning capability itself missing?
Building a minimal RLVR pipeline from scratch, understanding sequence log-probabilities, reward normalization, and training a reasoning model using exact-match rewards.
Let's say we have 4 GPUs on a single machine and we want to communicate between them — data is sliced, each GPU gets a different slice of the dataset, and the communication pattern differs depending on the type of model parallelism.
We are able to reduce the training time by almost 2.5x — which is quite significant when training large models for a longer time.
And that’s when we seriously began exploring distributed training — a shift that fundamentally changed how we build and scale our models.
20 July 2022
I hear a lot about self-attention in papers I read every now and then. And almost every time i get very puzzeled whenever the word is mentioned. Whether it is action recognition, image translation, image segementation or anything else (PS: sorry for being computer vision bias), self-attention is the new trend, I suppose.
25 December 2021
Fellowship.Ai is a four months long “unpaid” fellowship on various machine learning topics. The program is 100 % remote, and anyone can apply for it. (PS: Keep reading, if you should or not!). You can apply directly by submitting your resume, but if you complete one of the challenges mentioned on the website - you tend to increase your chances for getting to the interviews. The challenges are focused on different machine learning topics such as image segmentation, Natural Language Processing (NLP), One-Shot learning etc.