Staff Writer
Published August 24, 2026 · Updated October 7, 2026last updated dates

Agents That Learn From YouTube
Someone in the wild told people to do this yesterday.
The tweet came from a Spanish account called PA13L0. Nine hundred likes, seventeen hundred bookmarks, growing fast. The text, translated: "Former Google engineer shows you how to create self-improving AI agents. Just transcribe this video with Grok and pass it to your Hermes so it learns jiu-jitsu."
He was right. The reason it works tells you something about how agents should learn.
The Unit of Learning Is a Skill, Not a Model Update
Most people in this space think agent improvement means one of two things. Fine-tune a model on new data, or shove the data into a vector store and call it RAG. Both share the same assumption: learning means injecting information into the model's weights or into its retrieval context.
This is wrong for agents. Agents do not need to know more facts. They need to know how to do more things. That is a different problem.
When you watch someone tie a knot or configure a database, you do not rewire your brain at the synaptic level and you do not search the whole internet for similar knots before every attempt. You extract the procedure. Maybe you write it down. Next time you face the same situation, you follow the steps.
Skills are that. A structured procedure you can load into an agent's context on the next session. No weight update. No vector retrieval pipeline. Just a markdown file with a numbered procedure, some prerequisites, and a list of pitfalls the video warned you about.
The tweet worked because Hermes has a skill system. Someone watched a jiu-jitsu video, got the transcript, and asked Hermes to turn it into a skill. The agent read the procedure, saved it as a SKILL.md, and now every future session has access to those techniques. Not stored in weights, not embedded in a vector index. Loaded as a structured document on startup.
This is a different approach from how most people think about agent learning. There is research backing it up.
What the Literature Already Shows
NVIDIA published Voyager in 2023. They dropped an LLM-powered agent into Minecraft and let it explore autonomously. The key design decision was not the reinforcement learning algorithm or the fine-tuning pipeline. It was the skill library. Voyager wrote executable code for each skill it discovered, saved it, and retrieved relevant skills when facing new tasks. The agent got 3.3 times more unique items and hit tech tree milestones 15 times faster than prior systems. The authors noted explicitly that the agent bypassed model parameter fine-tuning entirely. The learning lived in the library, not in the model.
Reflexion came from Noah Shinn and others around the same time. Same insight. Instead of updating weights, the agent reflected on its failures, wrote down what it learned, and stored that text in an episodic memory buffer. The next attempt used the reflection to make better decisions. Reflexion hit 91% pass@1 on HumanEval, beating GPT-4's 80% at the time. All through verbal reinforcement learning. No gradient updates. Just better procedures.
A more recent paper, Learning Hierarchical Procedural Memory for LLM Agents, described a system called MACLA that extracts reusable procedures from trajectories, tracks their reliability through Bayesian posteriors, and refines them by contrasting successes against failures. Same architecture: frozen model, external procedure store, retrieval at inference time.
The pattern is consistent. The model does the inference. The skill library holds the memory. You update the library, not the model.
How the Pipeline Actually Works
The video-to-skill pipeline we built today is three steps.
Grab the transcript from YouTube. YouTube provides captions for most instructional content, and the youtube-transcript-api library fetches them with timestamps in one call. If the video has no transcript, you stop there. YouTube does not provide transcripts for all content.
Process the transcript into a procedure. This is the real work. The transcript is raw dialogue, not a recipe. Someone says "okay so what you want to do is..." and then demonstrates something with their hands for thirty seconds. An LLM reads the full transcript, identifies the procedural content, extracts the numbered steps, identifies the prerequisites and pitfalls mentioned in passing, and writes it up as a structured SKILL.md file. The skill has a name, a one-line description, a when-to-use section, numbered steps, and a pitfalls section.
Save it to the agent's skill directory. Next session, the agent loads the skill on startup. It shows up in skills_list. It is available for every relevant task going forward.
The whole thing takes maybe two minutes for a typical ten-minute tutorial video. The bottleneck is not the pipeline. It is whether the video has enough procedural content to extract.
Why This Is a Content Engineering Problem
Here is what changes. If learning means updating weights, then improving your agent requires GPU time, training data curation, and ML engineering. That is expensive, slow, and opaque. You do not know what the model learned, and you cannot selectively remove a bad lesson.
If learning means extracting procedures, then improving your agent requires good transcripts, clear instructions, and well-structured markdown. That is a content engineering problem, not an ML problem. It costs nothing to run. The result is inspectable, editable, and reversible. You can read exactly what the agent learned. You can fix a wrong step without retraining anything. You can share the skill file with someone else and they get the same capability.
This maps directly to how people learn. You watch someone do something. You take notes. You try it yourself. You refine your notes. The model in your head does not need retraining. The procedure in your notebook gets better.
Where It Breaks
The pipeline only works when the content is procedural. A video essay about the nature of consciousness is not going to produce a useful skill. Neither is a product launch keynote or a review video. The pipeline extracts steps. If there are no steps, there is nothing to extract.
Videos without transcripts are a hard block. YouTube generates auto-captions for most English content, but not everything. Some languages get worse coverage. Some creators turn captions off. Without a transcript, the pipeline has nothing to work with.
Vague content is the most common failure mode. A ten-minute tutorial where the creator talks around the subject, demonstrates things without explaining them, or jumps between topics without structure. The transcript is there, but the procedure is not extractable. The pipeline can produce a skill, but it will be shallow. It might capture the right keywords without encoding the actual technique.
The skill also only loads on the next session. Hermes caches skills at session start, so a skill created mid-conversation is not available until you start a new session. That is a product limitation, not a conceptual one. It matters for the user experience but not for the architecture.
Write a Skill Today
The single most useful thing you can do for your agent today is not to train a LoRA or set up a RAG pipeline or build a vector ingestion workflow. It is to write a skill. A clear, numbered procedure for something your agent does repeatedly. The skill system loaded on startup is a better return on time than any of those.
The video-to-skill pipeline reduces the cost of skill creation to basically zero for anything with a YouTube tutorial. Watch a video, get a skill. That is the loop. It is the same loop PA13L0 described in his tweet. Transcribe the video. Pass it to Hermes. Your agent learns something new by the next session.
Your agent does not need fine-tuning. It needs better skills.
