Troubleshooting large, distributed systems during production outages can be a daunting challenge. When downtime is not an option, how do you steer your ship through the storm of technical difficulties?
Drawing from extensive experience trying to keep production up at one of South Africa's largest banks, this talk delves into the realm of resilience engineering. It will equip you to identify and reproduce issues and also to make your systems more resilient in the face of chaos.
We will explore an array of troubleshooting tools, such as Application Performance Monitoring tools, logging, heap and thread dumps, application metrics, profiling techniques, load testing, and briefly touch on chaos engineering.
Recognizing that prevention is better than cure, we will conclude with patterns and strategies to help you build resilient systems. Topics will include timeouts, connection reuse, circuit breakers, bulkheads, and fallback mechanisms.
=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=--=-=-=-=
About Renette
Renette is a technical lead at Entelect and hold BIS Multimedia and BSc (Hons) Computer Science degrees from the University of Pretoria. She works on backend Java and Kotlin projects, primarily using Spring Boot. She is passionate about resilience engineering and keeping production up in tricky circumstances and loves sharing knowledge and mentoring. In her spare time, she enjoys reading fantasy books and playing board games.
On this page of the site you can watch the video online The One Pattern Every Distributed System Needs with a duration of hours minute second in good quality, which was uploaded by the user DevConf 01 November 2024, share the link with friends and acquaintances, this video has already been watched 160 times on youtube and it was liked by 9 viewers. Enjoy your viewing!