Discussion about this post

User's avatar
Khaled Ahmed, PhD's avatar

The evaluation + deployment axis is what most agent courses skip, so it's great you're putting it at the center. In my own research the part that keeps biting production teams isn't task success rate; it's trajectory-level evaluation, specifically whether each tool call was the right choice given the state the agent was in, not just whether the final answer was correct. Decomposing runs into atomic claims at the step level is how I ended up making these failures visible. One request: when you get to the evaluation week, consider framing MCP and A2A from the eval angle too, since the interfaces between services are where silent drift tends to live. Really looking forward to this.

Juan Guillermo Loaiza Andrade's avatar

I currently work as a data scientist at a Colombian startup (with over a year of experience), but I want to take the next step and get a remote job in the United States or another country with better career opportunities.

My bachelor degree wasn't in engineering but in business, so despite being a data scientist in title I'm constantly looking for ways to bridge that gap by acquiring the most relevant skills in the market. I have some knowledge of AI since my company uses it for classification, sentiment analysis, and so on.

My expectations for this training are to have projects I can showcase in my portfolio, as well as acquire the skills that will bring me closer to my goal of working abroad (a former colleague achieved this, so why can't I?). Do you think my expectations are realistic?

7 more comments...

No posts

Ready for more?