Abstract
At Neon digital bank in Brazil, we strive to make revenue-impacting predictions based on customer behavior. Building a low latency and high availability distributed system that meets this requirement becomes especially challenging.
In this talk, I will present how Neon improved the reliability, transparency and quality of its credit decisions by taking advantage of machine-learning models running on Tensorflow Serving and how we integrated the process with a credit approval backend on Go. I'll cover:
- How to carefully roll out a new credit modeling system using dark launches and observability tools.
- How Tensorflow Serving simplifies the serving path for machine-learning models, despite a few quirks and limitations.
- How to meet latency and reliability requirements through network proximity while complying with regulatory constraints.
Topics
QCon New York 2023 is a three day conference for senior software engineers, architects and team leads. An international program committee of working engineers selects every session. Patterns and practices, not products and pitches.
Part of the track
Designing Modern Reliable Architectures Hosted by Silvia Esparrachiari Software Engineer @GoogleFrom the same track
Wednesday 14 June
10:35 Salon A-C Session Kafka How to Build a Reliable Kafka Data Processing Pipeline, Focusing on Contention, Uptime and Latency Lily Mara Engineering Manager @OneSignal Shifting workloads from synchronous to asynchronous can simplify the operational cost of high-throughput HTTP services. But understanding the evolution of performance metrics in the world of complex, high-concurrency, asynchronous distributed systems can be quite challenging. 11:50 Carroll Gardens Unconference Unconference: Designing Modern Reliable Architectures What is an unconference? An unconference is a participant-driven meeting. Attendees come together, bringing their challenges and relying on the experience and know-how of their peers for solutions. 13:40 Salon A-C Session Architecture Building an Architecture to Predict Customer Behavior in a Revenue-Critical System Yves Junqueira Distinguished Software Engineer @Neon At Neon digital bank in Brazil, we strive to make revenue-impacting predictions based on customer behavior. Building a low latency and high availability distributed system that meets this requirement becomes especially challenging. 14:55 Salon A-C Session Architecture Reliable Architectures Through Observability Kent Quirk Staff Engineer @Honeycomb.io We want our systems to be reliable, but testing alone isn't enough. In a complex, multi-service system, it's impossible to test your way to correctness. That's why we need observability. Observability is the ability to see what our code is doing, in production and in development. 16:10 Salon A-C Session Developer Environment Architecting a Production Development Environment for Reliability Henrique Andrade Production Engineer @Meta At Meta, developers use a combination of development servers, including virtual machines and physical hosts, as well as on-demand containers to perform their daily software engineering work. 17:25 Salon A-C Session Cloud Architecture Survival Strategies for the Noisy Neighbor Apocalypse Meenakshi Jindal Staff Software Engineer @Netflix Noisy neighbor issues are a common challenge for multi-tenant platforms, leading to resource contention, performance degradation, and costly downtime for other tenants sharing the same resources.