As your code scales beyond a trivial level of complexity and sophistication it becomes difficult or impossible to know everything that it is doing. The flow of logic and data through your software and which parts are taking the most time are impossible to understand without help from your tools. VizTracer is the tool that you will turn to when you need to know all of the execution paths that are being exercised and which of those paths are the most expensive. In this episode Tian Gao explains why he created VizTracer and how you can use it to gain a deeper familiarity with the code that you are responsible for maintaining.
Summary
As your code scales beyond a trivial level of complexity and sophistication it becomes difficult or impossible to know everything that it is doing. The flow of logic and data through your software and which parts are taking the most time are impossible to understand without help from your tools. VizTracer is the tool that you will turn to when you need to know all of the execution paths that are being exercised and which of those paths are the most expensive. In this episode Tian Gao explains why he created VizTracer and how you can use it to gain a deeper familiarity with the code that you are responsible for maintaining.
Announcements
- Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
- When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. And now you can launch a managed MySQL, Postgres, or Mongo database cluster in minutes to keep your critical data safe with automated backups and failover. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
- Need to automate your Python code in the cloud? Want to avoid the hassle of setting up and maintaining infrastructure? Shipyard is the premier orchestration platform built to help you quickly launch, monitor, and share python workflows in a matter of minutes with 0 changes to your code. Shipyard provides powerful features like webhooks, error-handling, monitoring, automatic containerization, syncing with Github, and more. Plus, it comes with over 70 open-source, low-code templates to help you quickly build solutions with the tools you already use. Go to dataengineeringpodcast.com/shipyard to get started automating with a free developer plan today!
- Your host as usual is Tobias Macey and today I’m interviewing Tian Gao about VizTracer, a low-overhead logging/debugging/profiling tool that can trace and visualize your python code execution
Interview
- Introductions
- How did you get introduced to Python?
- Can you describe what VizTracer is and the story behind it?
- What are the main goals that you are focused on with VizTracer?
- What are some examples of the types of bugs that profiling can help diagnose?
- How does profiling work together with other debugging approaches? (e.g. logging, breakpoint debugging, etc.)
- There are a number of profiling utilities for Python. What feature or combination of features were missing that motivated you to create VizTracer?
- Can you describe how VizTracer is implemented?
- How have the design and goals changed since you started working on it?
- There are a number of styles of profiling, what was your process for deciding which approach to use?
- What are the most complex engineering tasks involved in building a profiling utility?
- Can you describe the process of using VizTracer to identify and debug errors and performance issues in a project?
- What are the options for using VizTracer in a production environment?
- What are the interfaces and extension points that you have built in to allow developers to customize VizTracer?
- What are some of the ways that you have used VizTracer while working on VizTracer?
- What are the most interesting, innovative, or unexpected ways that you have seen VizTracer used?
- What are the most interesting, unexpected, or challenging lessons that you have learned while working on VizTracer?
- When is VizTracer the wrong choice?
- What do you have planned for the future of VizTracer?
Keep In Touch
- gaogaotiantian on GitHub
Picks
- Tobias
- Travelers show on Netflix
- Tian
- objprint
- Lincoln Lawyer
- bilibili– Tian’s coding sessions in Chinese
Closing Announcements
- Thank you for listening! Don’t forget to check out our other shows. The Data Engineering Podcast covers the latest on modern data management. The Machine Learning Podcast helps you go from idea to production with machine learning.
- Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
- If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
- To help other people find the show please leave a review on iTunes and tell your friends and co-workers
Links
- Viztracer
- Python cProfile
- Sampling Profiler
- Perfetto
- Coverage.py
- Python setxprofile hook
- Circular Buffer
- Catapult Trace Viewer
- py-spy
- psutil
- gdb
- Flame graph
The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA