<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Systems in Practice]]></title><description><![CDATA[Real-world lessons in distributed systems, Java, Kafka, AWS and financial technology.]]></description><link>https://discere.tech</link><image><url>https://cdn.hashnode.com/uploads/logos/6a91a874a07da3ac2b00e1f2/ab893486-c5dc-4016-89d4-11ac03e9f01d.png</url><title>Systems in Practice</title><link>https://discere.tech</link></image><generator>RSS for Node</generator><lastBuildDate>Mon, 31 Aug 2026 01:40:22 GMT</lastBuildDate><atom:link href="https://discere.tech/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Leadership, Autonomy, and a Low-Hierarchy Working Culture]]></title><description><![CDATA[A reflections on guidance, responsibility, and trust in high-performing teams

If you need a team lead to show you a path, that can be useful and even enjoyable. If you need a manager to make sure you]]></description><link>https://discere.tech/leadership-autonomy-and-a-low-hierarchy-working-culture</link><guid isPermaLink="true">https://discere.tech/leadership-autonomy-and-a-low-hierarchy-working-culture</guid><dc:creator><![CDATA[Tarun Bansal]]></dc:creator><pubDate>Sun, 30 Aug 2026 20:40:13 GMT</pubDate><content:encoded><![CDATA[<p><em>A reflections on guidance, responsibility, and trust in high-performing teams</em></p>
<blockquote>
<p>If you need a team lead to show you a path, that can be useful and even enjoyable. If you need a manager to make sure you do your work, however, something is wrong on both sides.</p>
</blockquote>
<h2>Leadership is direction, not surveillance</h2>
<p>A strong team lead helps people see the landscape. They clarify the purpose, expose constraints, connect decisions to the wider system, and help the team discover a credible path forward.</p>
<p>This is leadership as navigation: the lead may know the terrain better, but the team still has to think, choose, and walk the path themselves.</p>
<p>That is very different from management as constant supervision. When an experienced professional needs someone to repeatedly check whether the work has started, whether promises will be kept, or whether quality matters, the problem is no longer a lack of guidance. It is a breakdown of ownership.</p>
<p>At the same time, when a manager assumes that capable people cannot be trusted without continuous oversight, the manager also contributes to that breakdown.</p>
<h2>What I learned from working in the Netherlands</h2>
<p>I am not a native Dutch person, so I cannot define a single "Dutch mindset." I can only describe the working habits I encountered and came to value while working in the Netherlands.</p>
<p>Many Dutch professional environments are associated with direct communication, relatively low hierarchy, and a willingness to challenge ideas regardless of job title. These qualities are not unique to the Netherlands, and not every Dutch workplace operates in the same way.</p>
<p>The healthiest version of this mindset is not bluntness for its own sake. It is the belief that adults at work should be able to speak plainly, explain their reasoning, question a decision, and accept responsibility for the outcome.</p>
<p>In such an environment, a title does not make an argument correct. A team lead earns influence by providing context, judgment, and consistency. Team members earn autonomy by being dependable, transparent, and willing to surface problems early.</p>
<p>Directness works only when it travels in both directions and remains respectful.</p>
<h2>What trust can look like in a Dutch workplace</h2>
<p>Trust at work is not merely a pleasant value. It changes how a team operates.</p>
<p>Here are some practical examples I associate with healthy, low-hierarchy teams in the Netherlands:</p>
<ul>
<li><p>A junior engineer can question a senior architect's proposal during a design discussion. The response should address the argument rather than invoke seniority.</p>
</li>
<li><p>A team member can say, "I disagree because this creates an operational risk," without that disagreement being treated as disloyalty.</p>
</li>
<li><p>People are generally expected to manage their own working day. The emphasis is on commitments, decisions, and outcomes rather than on appearing continuously busy.</p>
</li>
<li><p>A colleague working four days a week is not automatically considered less committed. The team plans around the person's agreed availability.</p>
</li>
<li><p>When someone has a family commitment or needs flexibility, the default response can be coordination rather than suspicion.</p>
</li>
<li><p>A team lead explains the goal and constraints, then allows the people closest to the problem to decide how to implement the solution.</p>
</li>
<li><p>Bad news is expected early. Saying that a deadline is at risk is viewed as responsible when it gives the team time to respond.</p>
</li>
</ul>
<p>This kind of trust is not blind. It depends on people doing what they promised, communicating when circumstances change, and accepting feedback directly.</p>
<h2>Consensus does not have to mean hierarchy</h2>
<p>The Netherlands also has a long tradition of consultation and negotiation, often described as the <strong>polder model</strong>. A Dutch government publication describes this as a tradition in which civil-society organizations, employers' associations, and trade unions were consulted before decisions were made.</p>
<p>That national model is not identical to how a software team works, but the principle is recognizable: involve the people affected by a decision, allow disagreement to become visible, and build enough shared understanding to move forward.</p>
<p>Consensus should not mean that every person has a veto or that decisions take forever. A strong lead listens broadly, makes the decision boundary clear, and closes the discussion when action is required.</p>
<h2>Fewer hours do not automatically mean less prosperity</h2>
<p>One reason this working culture interests me is that the Netherlands challenges the assumption that prosperity must come from spending the longest possible time at work.</p>
<p>According to <a href="https://ec.europa.eu/eurostat/web/products-eurostat-news/w/ddn-20260527-1">Eurostat's 2025 figures</a>, employed people in the Netherlands recorded the EU's shortest average actual working week in their main job: <strong>31.9 hours</strong>, compared with an EU average of roughly 36 hours. The national average is strongly influenced by the Netherlands' high level of part-time employment, so it should not be interpreted as every Dutch full-time employee working only 31.9 hours.</p>
<p>At the same time, <a href="https://ec.europa.eu/eurostat/statistics-explained/index.php?title=Purchasing_power_parities_and_GDP_per_capita_-_preliminary_estimate">Eurostat's preliminary 2025 comparison</a> placed Dutch GDP per capita, adjusted for price differences, <strong>more than 20% above the EU average</strong>. The <a href="https://www.oecd.org/en/publications/oecd-economic-surveys-netherlands-2025_2dd1f4aa-en/full-report/preserving-trade-competitiveness-amidst-increasing-global-fragmentation_d7457f46.html">OECD describes the Netherlands</a> as having a high standard of living, even while warning that recent productivity growth has slowed.</p>
<p>These figures do not prove that shorter hours create prosperity. Trade, infrastructure, education, capital, technology, institutions, labour-force participation, and many other factors contribute to economic performance. But they do challenge a simplistic belief:</p>
<blockquote>
<p>More hours at work do not automatically produce more value.</p>
</blockquote>
<p>The more useful question is what happens during those hours. Clear priorities, professional autonomy, good tools, candid communication, and mutual trust can allow people to create substantial value without turning visible exhaustion into a measure of commitment.</p>
<p>There is also statistical evidence that autonomy is not merely an impression. <a href="https://www.cbs.nl/en-gb/dossier/well-being-and-the-sustainable-development-goals/monitor-of-well-being-and-the-sustainable-development-goals-2023/the-sustainable-development-goals-in-the-monitor-of-well-being/sdg-s/sdg-8-2-labour-and-leisure-time">Statistics Netherlands reported</a> that <strong>65.9% of employed people aged 15-74 were free to decide how to do their work in 2022</strong>.</p>
<p>For me, that is the more interesting connection between culture and performance: trust is expressed through autonomy, but autonomy is sustained through responsibility.</p>
<h2>What a team lead should provide</h2>
<ul>
<li><p><strong>Purpose:</strong> Explain why the work matters and what outcome the team is trying to achieve.</p>
</li>
<li><p><strong>Context:</strong> Expose dependencies, constraints, risks, and organizational realities that are not visible from an individual task.</p>
</li>
<li><p><strong>A path when needed:</strong> Offer experience, options, and a starting direction without removing the team's responsibility to think.</p>
</li>
<li><p><strong>Decision clarity:</strong> Make or escalate decisions when the team cannot resolve ambiguity within the available time.</p>
</li>
<li><p><strong>Safety for candour:</strong> Make it acceptable to disagree, report bad news, and admit mistakes without unnecessary punishment.</p>
</li>
<li><p><strong>Coaching:</strong> Help people improve their judgment so that they need less direction over time, not more.</p>
</li>
</ul>
<h2>What every team member should provide</h2>
<ul>
<li><p><strong>Ownership:</strong> Understand the expected outcome and drive the work without waiting for repeated reminders.</p>
</li>
<li><p><strong>Visibility:</strong> Communicate progress, uncertainty, delays, and risks before they become surprises.</p>
</li>
<li><p><strong>Independent judgment:</strong> Bring possible solutions, not only problems, while remaining open to correction.</p>
</li>
<li><p><strong>Reliability:</strong> Honour commitments or renegotiate them early when the facts change.</p>
</li>
<li><p><strong>Constructive dissent:</strong> Challenge decisions with evidence and then support the agreed direction.</p>
</li>
<li><p><strong>Accountability:</strong> Own mistakes, repair their effects, and improve the system that allowed them.</p>
</li>
</ul>
<h2>Autonomy is reciprocal</h2>
<p>Autonomy is not the absence of accountability. It is a relationship built on evidence.</p>
<p>A manager gives people room to operate; team members make their work and reasoning visible enough that trust is rational. Neither side should require constant reassurance.</p>
<p>A useful test is whether the team can operate when the lead is unavailable for a few days.</p>
<p>If work stops because every decision requires permission, leadership has created dependency. If commitments disappear because nobody is watching, the team has confused freedom with a lack of responsibility.</p>
<p>A mature team avoids both extremes.</p>
<h2>Direct communication instead of unnecessary control</h2>
<p>Candid communication should replace much of the need for monitoring.</p>
<p>A team member should be able to say:</p>
<blockquote>
<p>I committed to Friday, but the integration risk is larger than I understood. I can deliver safely on Tuesday, or we can reduce scope and keep Friday.</p>
</blockquote>
<p>That statement gives the team facts, choices, and ownership. It is far more useful than a green status that turns red at the last moment.</p>
<p>A lead should respond with the same clarity:</p>
<blockquote>
<p>Tuesday affects another team, so let us reduce scope and preserve the safety work.</p>
</blockquote>
<p>The lead does not need to hover over implementation. They need to make the trade-off explicit and ensure that the decision is understood.</p>
<h2>When intervention is necessary</h2>
<p>Low hierarchy does not mean the absence of management.</p>
<p>A lead must intervene when risk is hidden, commitments are repeatedly missed, quality is knowingly compromised, or behaviour harms the team. The difference lies in the purpose of intervention: restore clarity and accountability rather than normalize permanent supervision.</p>
<p>The first response should normally be a direct conversation about expectations, evidence, and consequences.</p>
<ul>
<li><p>If the problem is missing skill, provide coaching.</p>
</li>
<li><p>If the problem is conflicting priorities, resolve them.</p>
</li>
<li><p>If expectations are unclear, make them explicit.</p>
</li>
<li><p>If ownership remains absent after expectations and support are clear, address it as a performance issue.</p>
</li>
</ul>
<p>One person's performance problem is not a reason to micromanage the entire team.</p>
<h2>A practical operating model</h2>
<ol>
<li><p>Agree on outcomes and boundaries, not a list of movements to monitor.</p>
</li>
<li><p>Make ownership explicit: one accountable person for each meaningful outcome.</p>
</li>
<li><p>Use short, factual progress updates focused on changes, risks, and decisions needed.</p>
</li>
<li><p>Escalate early when a commitment is threatened; do not wait for a status meeting.</p>
</li>
<li><p>Record important decisions and the reasoning behind them.</p>
</li>
<li><p>Review outcomes and learning, not hours of visible activity.</p>
</li>
<li><p>Increase autonomy when judgment and reliability grow; add support when either is missing.</p>
</li>
</ol>
<h2>The balanced principle</h2>
<p>A good lead can show a path, remove obstacles, and strengthen judgment. A good professional does not need a manager to manufacture motivation or enforce every ordinary responsibility.</p>
<p>High-performing teams emerge when leadership offers context and courage, while individuals offer ownership and candour.</p>
<blockquote>
<p><strong>A team lead should help you find the path; they should not have to push you along it.</strong></p>
</blockquote>
]]></content:encoded></item><item><title><![CDATA[System Design Interview Guide: A 5-Step Roadmap]]></title><description><![CDATA[System design interviews are not primarily tests of how many technologies you can name. They test whether you can turn unclear requirements into a sensible architecture, explain your decisions and rec]]></description><link>https://discere.tech/system-design-interview-guide-a-5-step-roadmap</link><guid isPermaLink="true">https://discere.tech/system-design-interview-guide-a-5-step-roadmap</guid><dc:creator><![CDATA[Tarun Bansal]]></dc:creator><pubDate>Fri, 28 Aug 2026 20:53:03 GMT</pubDate><content:encoded><![CDATA[<p>System design interviews are not primarily tests of how many technologies you can name. They test whether you can turn unclear requirements into a sensible architecture, explain your decisions and recognise the trade-offs behind them.</p>
<p>This guide organises the preparation into five steps:</p>
<pre><code class="language-mermaid">flowchart LR
    A["1. Fundamentals"] --&gt; B["2. Scalability"]
    B --&gt; C["3. Architecture patterns"]
    C --&gt; D["4. Data management"]
    D --&gt; E["5. Hands-on practice"]
</code></pre>
<p>The goal is not to memorise one perfect architecture. It is to build a repeatable way of thinking.</p>
<h2>Step 1: Fundamentals</h2>
<p>Before discussing distributed databases or event-driven systems, you need a strong understanding of how requests, APIs, storage and caching work.</p>
<h3>Networking</h3>
<h4>HTTP/1.1 versus HTTP/2</h4>
<p>HTTP/1.1 normally processes one outstanding request at a time per connection unless pipelining is used. Browsers often open several connections to compensate, but this adds overhead.</p>
<p>HTTP/2 uses binary framing and multiplexes many request and response streams over a single TCP connection. This makes it more efficient when a page or application needs many small resources.</p>
<h4>TCP and connection setup</h4>
<p>The TCP handshake uses three messages:</p>
<pre><code class="language-text">Client → SYN
Server → SYN-ACK
Client → ACK
</code></pre>
<p>This adds a network round trip before application data moves. At global scale, connection establishment becomes a meaningful part of latency. QUIC and HTTP/3 reduce some of this cost by combining transport security and connection establishment more efficiently.</p>
<h4>DNS</h4>
<p>DNS translates a domain name into an IP address. Results are cached at several levels:</p>
<ul>
<li><p>browser;</p>
</li>
<li><p>operating system;</p>
</li>
<li><p>local or corporate resolver;</p>
</li>
<li><p>recursive DNS resolver.</p>
</li>
</ul>
<p>Caching avoids repeating the complete lookup for every request, but it also means DNS changes take time to propagate according to the configured TTL.</p>
<h4>Layer 4 versus Layer 7 load balancing</h4>
<table>
<thead>
<tr>
<th>Layer</th>
<th>Uses</th>
<th>Strength</th>
<th>Trade-off</th>
</tr>
</thead>
<tbody><tr>
<td>L4</td>
<td>IP addresses, ports and transport information</td>
<td>Fast and protocol-agnostic</td>
<td>Cannot route using HTTP content</td>
</tr>
<tr>
<td>L7</td>
<td>Hostnames, paths, headers and cookies</td>
<td>Intelligent application routing</td>
<td>More processing and configuration</td>
</tr>
</tbody></table>
<p>An L7 load balancer can route <code>/images</code> and <code>/api</code> to different services, while an L4 load balancer forwards connections without understanding HTTP semantics.</p>
<h3>API design</h3>
<h4>REST versus GraphQL</h4>
<table>
<thead>
<tr>
<th>REST</th>
<th>GraphQL</th>
</tr>
</thead>
<tbody><tr>
<td>Models resources through endpoints</td>
<td>Models data through a query schema</td>
</tr>
<tr>
<td>Simple HTTP caching</td>
<td>More complicated caching</td>
</tr>
<tr>
<td>Can over-fetch or under-fetch</td>
<td>Clients request specific fields</td>
</tr>
<tr>
<td>Straightforward operational model</td>
<td>Flexible for clients with different data needs</td>
</tr>
<tr>
<td>Many endpoints may be required</td>
<td>One endpoint may serve many query shapes</td>
</tr>
</tbody></table>
<p>GraphQL flexibility also creates additional concerns around query cost, authorisation and abuse prevention.</p>
<h4>Rate limiting</h4>
<p>Rate limiting protects a service from overload and prevents one client from consuming all available capacity.</p>
<p><strong>Token bucket</strong> accumulates tokens at a steady rate. A request consumes a token, allowing short bursts while enforcing a long-term average.</p>
<p><strong>Sliding window</strong> counts requests within a continuously moving time interval. It provides smoother enforcement but typically requires more state.</p>
<h4>Sessions versus JWTs</h4>
<table>
<thead>
<tr>
<th>Server-side session</th>
<th>JWT</th>
</tr>
</thead>
<tbody><tr>
<td>State is stored by the server</td>
<td>Claims are carried in the token</td>
</tr>
<tr>
<td>Easy to revoke centrally</td>
<td>Difficult to revoke before expiry</td>
</tr>
<tr>
<td>Requires shared session storage when scaling</td>
<td>Easy for multiple service instances to verify</td>
</tr>
<tr>
<td>Session ID is normally stored in a cookie</td>
<td>Token must be protected from theft and misuse</td>
</tr>
</tbody></table>
<p>JWT revocation is difficult precisely because a service can validate the token without consulting central state. Common mitigations include short expiries, refresh-token rotation and revocation lists.</p>
<h3>Database basics</h3>
<h4>SQL versus NoSQL</h4>
<p>Relational databases provide schemas, joins, constraints and transactions. NoSQL systems often provide flexible models and easier horizontal distribution for specific access patterns, but usually trade away some relational capabilities.</p>
<p>Neither category is automatically more scalable. The right choice depends on queries, consistency requirements, data shape and operating scale.</p>
<h4>B-tree versus hash indexes</h4>
<table>
<thead>
<tr>
<th>B-tree index</th>
<th>Hash index</th>
</tr>
</thead>
<tbody><tr>
<td>Maintains sorted keys</td>
<td>Maps keys through a hash function</td>
</tr>
<tr>
<td>Supports exact matches</td>
<td>Excellent for exact matches</td>
</tr>
<tr>
<td>Supports range and ordered queries</td>
<td>Does not naturally support range scans</td>
</tr>
<tr>
<td>Useful for dates and sortable values</td>
<td>Useful for equality-based lookups</td>
</tr>
</tbody></table>
<h4>Range versus hash partitioning</h4>
<p>Range partitioning groups nearby key values together. It supports efficient range queries but may create hot partitions when activity concentrates in one range.</p>
<p>Hash partitioning spreads keys more evenly, improving load distribution while reducing range-query locality.</p>
<h3>Caching</h3>
<table>
<thead>
<tr>
<th>Pattern</th>
<th>How it works</th>
<th>Main trade-off</th>
</tr>
</thead>
<tbody><tr>
<td>Cache-aside</td>
<td>Application reads the cache and loads from the database on a miss</td>
<td>Simple, but cached data can become stale</td>
</tr>
<tr>
<td>Write-through</td>
<td>Writes update the cache and durable store together</td>
<td>Consistent reads, slower writes</td>
</tr>
<tr>
<td>Write-behind</td>
<td>Writes enter the cache and reach storage asynchronously</td>
<td>Fast writes, risk during failures</td>
</tr>
</tbody></table>
<p>A <strong>cache stampede</strong> occurs when a popular entry expires and many requests simultaneously query the database. Mitigations include:</p>
<ul>
<li><p>per-key locks;</p>
</li>
<li><p>request coalescing;</p>
</li>
<li><p>randomised or staggered expiry;</p>
</li>
<li><p>refreshing popular values before expiry;</p>
</li>
<li><p>serving stale values briefly while refreshing.</p>
</li>
</ul>
<p>A CDN caches content near users at edge locations. It is particularly effective for static files, images and video. An application or origin cache normally sits closer to the service and can cache dynamic or computed results.</p>
<h2>Step 2: Scalability and Performance</h2>
<h3>Vertical versus horizontal scaling</h3>
<table>
<thead>
<tr>
<th>Vertical scaling</th>
<th>Horizontal scaling</th>
</tr>
</thead>
<tbody><tr>
<td>Add CPU, memory or storage to one machine</td>
<td>Add more machines</td>
</tr>
<tr>
<td>Operationally simple</td>
<td>Scales beyond one machine's limit</td>
</tr>
<tr>
<td>Eventually reaches a hardware ceiling</td>
<td>Requires load distribution and coordination</td>
</tr>
<tr>
<td>Can remain a single point of failure</td>
<td>Introduces distributed-system failures</td>
</tr>
</tbody></table>
<p>Horizontal scaling is not free. Function calls become network calls, replicas can disagree and failures may be partial rather than complete.</p>
<h3>Load-balancing algorithms</h3>
<p><strong>Round robin</strong> cycles evenly through servers. It is simple, but assumes requests and servers have similar cost and performance.</p>
<p><strong>Least connections</strong> sends a request to the server with the fewest active connections. It adapts better to unequal request duration, although connection count is not always a perfect measure of actual load.</p>
<p><strong>Consistent hashing</strong> places servers and keys on a hash ring. Adding or removing a server moves only part of the keyspace. This is useful for distributed caches because it preserves cache locality during membership changes.</p>
<h3>Replication</h3>
<h4>Leader-follower</h4>
<p>Writes go to one leader and are copied to followers. Followers can serve reads, but replication lag means a user may not immediately observe a completed write when reading from a follower.</p>
<h4>Multi-leader</h4>
<p>Several nodes accept writes. This is useful across regions or disconnected environments, but simultaneous changes can conflict and require a resolution policy.</p>
<h4>Quorums</h4>
<p>For <code>N</code> replicas, a write may require acknowledgements from <code>W</code> replicas and a read may consult <code>R</code> replicas. Increasing the quorum improves consistency confidence but adds latency and can reduce availability during failures.</p>
<h3>Sharding</h3>
<table>
<thead>
<tr>
<th>Strategy</th>
<th>Advantage</th>
<th>Risk</th>
</tr>
</thead>
<tbody><tr>
<td>Range sharding</td>
<td>Efficient range queries</td>
<td>Hot ranges and uneven growth</td>
</tr>
<tr>
<td>Hash sharding</td>
<td>Even distribution</td>
<td>Poor range-query locality</td>
</tr>
<tr>
<td>Directory-based sharding</td>
<td>Flexible placement</td>
<td>Extra lookup and directory dependency</td>
</tr>
</tbody></table>
<p>Resharding is the operational challenge of redistributing data when shards are added, removed or become unbalanced. Consistent hashing and virtual nodes reduce movement, but migration, dual writes and cutover still require careful design.</p>
<h3>Asynchronous processing</h3>
<h4>Kafka versus RabbitMQ</h4>
<table>
<thead>
<tr>
<th>Kafka</th>
<th>RabbitMQ</th>
</tr>
</thead>
<tbody><tr>
<td>Distributed append-only log</td>
<td>Traditional message broker</td>
</tr>
<tr>
<td>High-throughput streaming</td>
<td>Flexible message routing</td>
</tr>
<tr>
<td>Consumers track positions and can replay</td>
<td>Per-message acknowledgement and queues</td>
</tr>
<tr>
<td>Strong fit for event streams and analytics</td>
<td>Strong fit for work queues and routing patterns</td>
</tr>
</tbody></table>
<p>This is not an absolute rule. The choice depends on retention, replay, ordering, routing and operating model.</p>
<h4>Delivery guarantees</h4>
<p>At-least-once delivery means a message should not be lost, but it may be processed more than once. Consumers must therefore be idempotent or deduplicate using a stable event ID.</p>
<p>Exactly-once delivery is much harder to guarantee end to end because messaging, application logic and database writes cross different boundaries. Practical designs usually combine broker guarantees with transactional writes, idempotency and deduplication.</p>
<h4>Backpressure</h4>
<p>Backpressure is how a system prevents fast producers from overwhelming slower consumers. Possible responses include:</p>
<ul>
<li><p>slowing or rejecting producers;</p>
</li>
<li><p>limiting concurrency;</p>
</li>
<li><p>pausing consumption;</p>
</li>
<li><p>buffering within bounded queues;</p>
</li>
<li><p>shedding low-priority work;</p>
</li>
<li><p>scaling consumers when additional parallelism is available.</p>
</li>
</ul>
<p>Without backpressure, queues grow without limit, latency rises and the system eventually exhausts resources.</p>
<h2>Step 3: Architecture Patterns</h2>
<h3>Monolith versus microservices</h3>
<table>
<thead>
<tr>
<th>Monolith</th>
<th>Microservices</th>
</tr>
</thead>
<tbody><tr>
<td>One deployable unit</td>
<td>Independently deployable services</td>
</tr>
<tr>
<td>Easier local development and testing</td>
<td>Independent ownership and scaling</td>
</tr>
<tr>
<td>Simple transactions and calls</td>
<td>Network boundaries and distributed data</td>
</tr>
<tr>
<td>Deployment coupling increases with size</td>
<td>Operational complexity increases with service count</td>
</tr>
</tbody></table>
<p>The real trade-off is often <strong>deployment coupling versus operational complexity</strong>. A well-designed monolith can scale effectively, while poorly bounded microservices can make every change harder.</p>
<h3>Event-driven versus request-response</h3>
<p>Request-response is synchronous. The caller waits for the result, which is easy to understand but couples its latency and availability to the downstream service.</p>
<p>Event-driven interaction is asynchronous. Producers publish facts and consumers react later. This decouples services in time, but introduces eventual consistency and more complicated failure diagnosis.</p>
<p>Eventual consistency is acceptable when the business can tolerate temporary staleness, such as a delayed social-media counter. It may be unacceptable when checking a bank balance immediately before a withdrawal.</p>
<h3>CQRS and event sourcing</h3>
<p><strong>CQRS</strong> separates the command model used for validating and persisting changes from the query model used for reads. It is valuable when reads and writes differ substantially in shape or scale.</p>
<p><strong>Event sourcing</strong> stores state changes as immutable events instead of overwriting the current record. Current state is reconstructed by replaying events.</p>
<p>Benefits include auditability, historical reconstruction and replay. Costs include event evolution, storage, debugging complexity and projection management.</p>
<p>Snapshots periodically store derived state so a service does not need to replay the complete history on every recovery.</p>
<h3>Fault tolerance</h3>
<h4>Circuit breaker</h4>
<p>A circuit breaker stops repeatedly calling an unhealthy dependency. It fails fast while the circuit is open and periodically probes whether the dependency has recovered.</p>
<h4>Retries with backoff and jitter</h4>
<p>Backoff increases the delay between attempts. Jitter adds randomness so many clients do not retry at exactly the same moment and overwhelm a recovering service.</p>
<p>Retries should be used only for transient and safe-to-repeat operations. Retrying an invalid request or a non-idempotent operation can make the problem worse.</p>
<h4>Bulkheads</h4>
<p>Bulkheads isolate resources such as thread pools, queues and connection pools per dependency. A failure in one integration then cannot consume every resource needed by the rest of the application.</p>
<h2>Step 4: Data Management</h2>
<h3>CAP theorem</h3>
<p>During a network partition, a distributed system must choose between:</p>
<ul>
<li><p><strong>Consistency:</strong> every read observes the required latest state;</p>
</li>
<li><p><strong>Availability:</strong> every request receives a non-error response, even when that response may be stale.</p>
</li>
</ul>
<p>Partition tolerance is not normally optional in a distributed system. CAP is therefore about behaviour during a partition, not about permanently labelling an entire product as only “CP” or “AP”. Different operations may choose different trade-offs.</p>
<h3>PACELC</h3>
<p>PACELC extends the discussion:</p>
<pre><code class="language-text">If there is a Partition: choose Availability or Consistency.
Else: choose Latency or Consistency.
</code></pre>
<p>Even during healthy operation, synchronously coordinating replicas improves consistency at the cost of latency.</p>
<h3>Consistency models</h3>
<table>
<thead>
<tr>
<th>Model</th>
<th>Guarantee</th>
</tr>
</thead>
<tbody><tr>
<td>Strong consistency</td>
<td>Reads observe the latest successful write according to the system's contract</td>
</tr>
<tr>
<td>Eventual consistency</td>
<td>Replicas converge, but a read may temporarily be stale</td>
</tr>
<tr>
<td>Causal consistency</td>
<td>Causally related operations are observed in order</td>
</tr>
<tr>
<td>Read-your-writes</td>
<td>A client observes its own completed writes</td>
</tr>
</tbody></table>
<p>Choose the model from the business invariant. A product catalogue and a payment balance do not necessarily need the same guarantee.</p>
<h3>Choosing a storage system</h3>
<p>A relational database with suitable indexes can handle substantial scale while preserving joins, constraints and transactions. Most systems should not abandon relational modelling merely because NoSQL sounds more scalable.</p>
<p>NoSQL becomes compelling when the workload has characteristics such as:</p>
<ul>
<li><p>extremely high write throughput;</p>
</li>
<li><p>flexible or rapidly changing document structures;</p>
</li>
<li><p>access patterns that do not benefit from joins;</p>
</li>
<li><p>very wide time-series data;</p>
</li>
<li><p>simple key-based access distributed across many partitions.</p>
</li>
</ul>
<p>The decision should begin with queries and invariants, not a technology label.</p>
<h3>Query optimisation</h3>
<p>An <code>EXPLAIN</code> plan shows how the database intends to execute a query, including scans, index access, join order and row estimates. It should be the first tool used when investigating a slow query.</p>
<p>A <strong>covering index</strong> contains every column required by a query, allowing the database to answer from the index without reading the full table row.</p>
<p>The <strong>N+1 query problem</strong> occurs when an application fetches a list and then runs one additional query for every item. It can often be fixed with a join, batch fetch or one <code>IN (...)</code> query.</p>
<h2>Step 5: Hands-On Practice</h2>
<h3>Design systems end to end</h3>
<p>Do not stop at a whiteboard diagram. Define enough detail to expose hidden assumptions:</p>
<ul>
<li><p>functional and non-functional requirements;</p>
</li>
<li><p>traffic and storage estimates;</p>
</li>
<li><p>API request and response shapes;</p>
</li>
<li><p>status codes and error contracts;</p>
</li>
<li><p>database schema and indexes;</p>
</li>
<li><p>partition keys;</p>
</li>
<li><p>pagination;</p>
</li>
<li><p>concurrency behaviour;</p>
</li>
<li><p>consistency boundaries;</p>
</li>
<li><p>failure and recovery paths;</p>
</li>
<li><p>monitoring and capacity signals.</p>
</li>
</ul>
<p>This is where vague designs become testable designs.</p>
<h3>Document trade-offs explicitly</h3>
<p>For each major decision, state:</p>
<pre><code class="language-text">We chose X because we are optimising for Y.
The cost is Z.
We would reconsider this decision if condition A changed.
</code></pre>
<p>Interviewers are usually more interested in whether you understand the consequences than whether you chose their favourite database.</p>
<h3>Practise out loud</h3>
<p>Speaking through a design under time pressure is different from silently writing one. Practise with a peer, mentor or recording. Train yourself to:</p>
<ol>
<li><p>clarify the problem;</p>
</li>
<li><p>state assumptions;</p>
</li>
<li><p>estimate scale;</p>
</li>
<li><p>draw the high-level architecture;</p>
</li>
<li><p>deepen the critical path;</p>
</li>
<li><p>identify bottlenecks and failures;</p>
</li>
<li><p>explain alternatives and trade-offs.</p>
</li>
</ol>
<h3>Debrief each mock interview</h3>
<p>After every mock, identify the most important question you should have asked earlier. Common omissions include:</p>
<ul>
<li><p>expected traffic;</p>
</li>
<li><p>read-to-write ratio;</p>
</li>
<li><p>latency target;</p>
</li>
<li><p>data-retention period;</p>
</li>
<li><p>regional distribution;</p>
</li>
<li><p>consistency requirement;</p>
</li>
<li><p>largest expected object;</p>
</li>
<li><p>acceptable data loss;</p>
</li>
<li><p>recovery-time objective.</p>
</li>
</ul>
<p>Designing for the wrong constraints is one of the most common system-design mistakes.</p>
<h2>A Repeatable Interview Framework</h2>
<p>Use this sequence during the interview:</p>
<h3>1. Clarify requirements</h3>
<p>Separate must-have functionality from optional scope. Ask about users, scale, latency, consistency, security and availability.</p>
<h3>2. Estimate scale</h3>
<p>Calculate rough requests per second, storage growth, bandwidth and read/write ratios. Approximate numbers are enough if the assumptions are explicit.</p>
<h3>3. Define APIs and data</h3>
<p>Identify core entities, API operations, schemas, indexes and partition keys.</p>
<h3>4. Draw the high-level flow</h3>
<p>Show clients, gateways, services, storage, caches and asynchronous components. Keep the first diagram simple.</p>
<h3>5. Deep-dive into the hardest path</h3>
<p>Choose the most important risk: fan-out, ordering, consistency, hot keys, search, delivery guarantees or multi-region failover.</p>
<h3>6. Cover failures and operations</h3>
<p>Explain retries, timeouts, idempotency, replication, monitoring, recovery and deployment.</p>
<h3>7. Summarise the trade-offs</h3>
<p>End by stating what the design optimises for, what it sacrifices and what would change at ten times the scale.</p>
<h2>Systems Worth Designing</h2>
<p>Practise these systems end to end:</p>
<ul>
<li><p>URL shortener;</p>
</li>
<li><p>distributed rate limiter;</p>
</li>
<li><p>news feed or timeline;</p>
</li>
<li><p>chat application;</p>
</li>
<li><p>distributed cache;</p>
</li>
<li><p>file-storage or Dropbox-like platform;</p>
</li>
<li><p>ride-sharing dispatch;</p>
</li>
<li><p>notification system.</p>
</li>
</ul>
<p>For each exercise, avoid copying a standard diagram. Change one important constraint and observe how the design changes. For example:</p>
<ul>
<li><p>require global active-active writes;</p>
</li>
<li><p>require strict ordering per user;</p>
</li>
<li><p>allow no data loss;</p>
</li>
<li><p>support very large files;</p>
</li>
<li><p>add a 100 ms latency target;</p>
</li>
<li><p>require data residency by country.</p>
</li>
</ul>
<h2>Recommended YouTube Resources</h2>
<ol>
<li><p><a href="https://lnkd.in/dBhRp73y">Shrayansh Jain</a></p>
</li>
<li><p><a href="https://lnkd.in/dXFamdZn">Rajat Gajbhiye</a></p>
</li>
<li><p><a href="https://lnkd.in/dN24R9n2">Gaurav Sen</a></p>
</li>
<li><p><a href="https://lnkd.in/ddP-WMaH">Arpit Bhayani</a></p>
</li>
</ol>
<h2>Final Advice</h2>
<p>A strong system design answer is not the biggest possible architecture. It is an architecture that matches the stated requirements and whose weaknesses you understand.</p>
<p>Start simple. Make assumptions explicit. Deepen the parts that carry the most risk. Explain what fails, how the system recovers and why each major trade-off is acceptable.</p>
<p>That is the skill the interview is really testing.</p>
]]></content:encoded></item><item><title><![CDATA[Concurrency]]></title><description><![CDATA[Thread Lifecycle
NEW → RUNNABLE → BLOCKED/WAITING/TIMED_WAITING → TERMINATED

BLOCKED: waiting to acquire a monitor lock

WAITING: Object.wait(), Thread.join() with no timeout — needs explicit notify
]]></description><link>https://discere.tech/concurrency</link><guid isPermaLink="true">https://discere.tech/concurrency</guid><dc:creator><![CDATA[Tarun Bansal]]></dc:creator><pubDate>Fri, 28 Aug 2026 20:50:07 GMT</pubDate><content:encoded><![CDATA[<hr />
<h3>Thread Lifecycle</h3>
<p><code>NEW → RUNNABLE → BLOCKED/WAITING/TIMED_WAITING → TERMINATED</code></p>
<ul>
<li><p><strong>BLOCKED</strong>: waiting to acquire a monitor lock</p>
</li>
<li><p><strong>WAITING</strong>: <code>Object.wait()</code>, <code>Thread.join()</code> with no timeout — needs explicit notify</p>
</li>
<li><p><strong>TIMED_WAITING</strong>: <code>sleep()</code>, <code>wait(timeout)</code>, <code>join(timeout)</code> — auto-wakes</p>
</li>
</ul>
<pre><code class="language-plaintext">        ┌─────┐
        │ NEW │
        └──┬──┘
           │ start()
           ▼
      ┌─────────┐   lock contention   ┌─────────┐
      │RUNNABLE │ ──────────────────▶ │ BLOCKED │
      │         │ ◀────────────────── │         │
      └────┬────┘   lock acquired     └─────────┘
           │
           │ wait()/join()            ┌─────────┐
           └────────────────────────▶ │ WAITING │
           ◀──────────── notify()──── └─────────┘
           │
           │ sleep(n)/wait(n)         ┌──────────────┐
           └────────────────────────▶ │TIMED_WAITING │
           ◀──────────── timeout ──── └──────────────┘
           │
           ▼
      ┌────────────┐
      │ TERMINATED │
      └────────────┘
</code></pre>
<h3>volatile vs synchronized</h3>
<p><code>volatile</code> guarantees visibility and happens-before, but <strong>not</strong> atomicity. <code>synchronized</code> guarantees both visibility and atomicity.</p>
<p><code>int++</code> is not atomic even on <code>volatile</code> — it's three operations: read, increment, write. Two threads can interleave.</p>
<pre><code class="language-java">// volatile — visibility only
volatile int counter = 0;
counter++;  // NOT ATOMIC! Read-Modify-Write can interleave

// Thread 1: read(0) → increment → write(1)
// Thread 2: read(0) → increment → write(1)  ← lost update!
// Result: 1 instead of 2

// Correct: use AtomicInteger
AtomicInteger counter = new AtomicInteger(0);
counter.incrementAndGet(); // CAS — truly atomic

// synchronized — atomicity + visibility
private int count = 0;
synchronized void increment() {
    count++;  // safe: only one thread at a time
}

// volatile IS sufficient for a single write/read (flag pattern)
volatile boolean shutdown = false;
// Thread A: shutdown = true;   (single write — atomic for boolean)
// Thread B: while (!shutdown)  (reads fresh value — visible)
</code></pre>
<h3>Key Synchronizers</h3>
<pre><code class="language-java">// ReentrantLock — explicit lock with tryLock, timed lock
ReentrantLock lock = new ReentrantLock();
lock.lock();
try { /* critical section */ }
finally { lock.unlock(); }

// ReadWriteLock — many readers OR one writer
ReadWriteLock rwLock = new ReentrantReadWriteLock();
rwLock.readLock().lock();  // multiple threads can hold simultaneously
rwLock.readLock().unlock();

// CountDownLatch — wait for N events (one-shot)
CountDownLatch latch = new CountDownLatch(3);
// 3 worker threads each call latch.countDown()
latch.await(); // main thread waits until count reaches 0
// Cannot be reset

// CyclicBarrier — N threads meet at a point, then continue together
CyclicBarrier barrier = new CyclicBarrier(3, () -&gt; System.out.println("All ready"));
// each thread calls barrier.await() — last one triggers the action
// Can be reset and reused

// Semaphore — limit concurrent access
Semaphore dbPool = new Semaphore(10); // max 10 concurrent DB connections
dbPool.acquire();
try { /* use connection */ }
finally { dbPool.release(); }
</code></pre>
<h3>Thread Pools — ThreadPoolExecutor &amp; ForkJoinPool</h3>
<p><code>ThreadPoolExecutor</code> internals: task submission checks <code>corePoolSize</code> → queue → <code>maximumPoolSize</code> → <code>RejectedExecutionHandler</code>. The queue type drives when new threads are created. <code>ForkJoinPool</code> uses work-stealing — idle threads steal tasks from busy threads' deques.</p>
<pre><code class="language-plaintext">Task submitted
       │
       ▼
  core threads &lt; corePoolSize?
       │ YES → create new thread
       │ NO
       ▼
  Queue full?
       │ NO → enqueue task
       │ YES
       ▼
  threads &lt; maxPoolSize?
       │ YES → create new thread
       │ NO
       ▼
  RejectedExecutionHandler
</code></pre>
<pre><code class="language-java">// ThreadPoolExecutor — explicit control
ExecutorService pool = new ThreadPoolExecutor(
    4,                              // corePoolSize
    8,                              // maximumPoolSize
    60, TimeUnit.SECONDS,           // keepAlive for extra threads
    new ArrayBlockingQueue&lt;&gt;(100),  // bounded queue
    new ThreadPoolExecutor.CallerRunsPolicy() // rejection: caller runs it
);

// ForkJoinPool — divide and conquer
ForkJoinPool fjp = new ForkJoinPool(4); // parallelism = 4
fjp.invoke(new RecursiveTask&lt;Integer&gt;() {
    protected Integer compute() {
        if (problem is small) return solve();
        // split into two sub-tasks
        var left  = new SubTask(leftHalf).fork();
        var right = new SubTask(rightHalf).fork();
        return left.join() + right.join();
    }
});

// Common factory methods (use carefully)
Executors.newFixedThreadPool(4);      // bounded threads, unbounded queue (!)
Executors.newCachedThreadPool();      // unbounded threads — danger at scale
Executors.newWorkStealingPool();      // wraps ForkJoinPool
</code></pre>
<h3>CompletableFuture</h3>
<p><code>CompletableFuture</code> enables non-blocking async pipelines. Key distinction: <code>thenApply</code> runs on the completing thread (sync), <code>thenApplyAsync</code> runs on a pool. <code>thenCompose</code> flattens nested futures.</p>
<pre><code class="language-java">// Basic async pipeline
CompletableFuture&lt;String&gt; future = CompletableFuture
    .supplyAsync(() -&gt; fetchUser(id))          // runs on ForkJoinPool
    .thenApply(user -&gt; user.getName())         // sync: same thread
    .thenApplyAsync(name -&gt; enrich(name))      // async: pool thread
    .thenCompose(name -&gt; fetchOrders(name));   // flatMap — avoids CF&lt;CF&lt;T&gt;&gt;

// Combining futures
CompletableFuture&lt;User&gt;   userFuture   = fetchUserAsync(id);
CompletableFuture&lt;Account&gt; accountFuture = fetchAccountAsync(id);

CompletableFuture.allOf(userFuture, accountFuture)
    .thenRun(() -&gt; {
        User user       = userFuture.join();
        Account account = accountFuture.join();
        combine(user, account);
    });

// Error handling
CompletableFuture&lt;String&gt; safe = future
    .exceptionally(ex -&gt; "fallback-value")    // recover from exception
    .handle((result, ex) -&gt; {                 // always runs
        if (ex != null) return "error";
        return result.toUpperCase();
    });

// Timeout (Java 9+)
future.orTimeout(5, TimeUnit.SECONDS)
      .exceptionally(ex -&gt; "timed out");
</code></pre>
<h3>The Most Dangerous Concurrency Bugs</h3>
<pre><code class="language-java">// DEADLOCK — always acquire locks in the same order
// Thread 1: lock(A) then lock(B)
// Thread 2: lock(B) then lock(A)  → circular wait
// Fix: enforce global lock ordering (e.g. by ID)
if (a.id &lt; b.id) { lock(a); lock(b); }
else             { lock(b); lock(a); }

// FALSE SHARING — threads write different fields on same cache line
class Counter {
    volatile long a;  // Thread 1 writes
    volatile long b;  // Thread 2 writes — same 64-byte cache line!
    // Each write invalidates the other thread's cache
}
// Fix: @Contended (JVM flag: -XX:-RestrictContended)
@jdk.internal.vm.annotation.Contended volatile long a;

// THREADLOCAL LEAK in thread pools
static ThreadLocal&lt;Connection&gt; conn = new ThreadLocal&lt;&gt;();
// Thread from pool: conn.set(c);
// Task ends but thread lives on → Connection never closed
// Fix: always conn.remove() in finally block

// LIVELOCK — threads keep responding to each other, no progress
// Thread A: see conflict → back off → retry → see conflict again
// Thread B: same pattern — neither makes progress
// Fix: randomized backoff
</code></pre>
]]></content:encoded></item><item><title><![CDATA[Java Garbage Collection Cheat Sheet]]></title><description><![CDATA[A practical reference covering GC collector trade-offs, G1's region-based layout, low-pause collectors (ZGC/Shenandoah), and how to diagnose GC behavior in production.

Collector Trade-offs
Trade-off ]]></description><link>https://discere.tech/java-garbage-collection-cheat-sheet</link><guid isPermaLink="true">https://discere.tech/java-garbage-collection-cheat-sheet</guid><dc:creator><![CDATA[Tarun Bansal]]></dc:creator><pubDate>Fri, 28 Aug 2026 20:33:02 GMT</pubDate><content:encoded><![CDATA[<p>A practical reference covering GC collector trade-offs, G1's region-based layout, low-pause collectors (ZGC/Shenandoah), and how to diagnose GC behavior in production.</p>
<hr />
<h2>Collector Trade-offs</h2>
<p>Trade-off triangle: throughput vs pause time vs memory footprint. No GC wins on all three.</p>
<pre><code class="language-plaintext">Collector      Throughput   Pause Time   Use Case
─────────────────────────────────────────────────────
Serial         Low          High         Single-core, small heaps
Parallel       High         Medium       Batch processing, throughput focus
CMS            Medium       Low(ish)     Deprecated in Java 14
G1 (default)   High         Predictable  General purpose (Java 9+)
ZGC            High         Sub-ms       Latency-critical (Java 15+)
Shenandoah     High         Sub-ms       Same as ZGC, different algorithm
</code></pre>
<hr />
<h2>G1 — Region-Based Heap Layout</h2>
<p>Heap divided into ~2048 equal-sized regions (1–32 MB each). Regions can be Eden, Survivor, Old, or Humongous (large objects). G1 prioritizes regions with most garbage first — hence "Garbage First". Mixed GC collects both young and old regions together.</p>
<hr />
<h2>ZGC &amp; Shenandoah — Concurrent Low-Pause Collectors</h2>
<p>Both achieve sub-millisecond pauses by doing most GC work concurrently with the application. They use load barriers to handle object references being moved while the app runs. Higher CPU overhead (~5–15%) — the cost of concurrency. Ideal for payment APIs and low-latency services.</p>
<pre><code class="language-bash"># Enable ZGC (Java 15+ for production)
-XX:+UseZGC
-XX:SoftMaxHeapSize=4g         # soft limit — ZGC tries to stay under this

# ZGC phases (all concurrent, no stop-the-world except tiny pauses):
# 1. Mark start (pause ~1ms)
# 2. Concurrent mark
# 3. Mark end (pause ~1ms)
# 4. Concurrent process references
# 5. Concurrent relocate
# 6. Concurrent remap

# Shenandoah
-XX:+UseShenandoahGC
-XX:ShenandoahGCMode=iu       # incremental-update mode (default)
</code></pre>
<hr />
<h2>GC Logging &amp; Diagnostics</h2>
<p>Enable GC logging with <code>-Xlog:gc*</code> (Java 9+). Key things to look for: frequency of collections, pause duration, heap size before/after, allocation rate. A Full GC is always a red flag — it stops all threads.</p>
<pre><code class="language-bash"># Enable GC logging
-Xlog:gc*:file=gc.log:time,uptime,level,tags

# Sample G1 log output:
[2.456s][info][gc] GC(3) Pause Young (Normal) (G1 Evacuation Pause)
[2.456s][info][gc] GC(3)   Heap: 512M -&gt; 128M (1024M)
[2.456s][info][gc] GC(3)   Pause: 12.3ms

# Red flags:
# - "Full GC" → heap pressure, possible leak, or wrong GC tuning
# - Pause &gt; MaxGCPauseMillis target consistently
# - Heap size after GC growing each cycle → leak
# - Allocation rate spiking → short-lived object pressure

# Useful tools:
# - GCViewer (open source)
# - GCEasy (web-based)
# - JDK's built-in: jstat -gcutil &lt;pid&gt; 1000
</code></pre>
]]></content:encoded></item><item><title><![CDATA[Java Memory Model & Concurrency Cheat Sheet]]></title><description><![CDATA[A practical reference covering JVM memory internals and Java concurrency — heap structure, the Java Memory Model, thread lifecycle, synchronization primitives, executors, CompletableFuture, and the co]]></description><link>https://discere.tech/java-memory-model-concurrency-cheat-sheet</link><guid isPermaLink="true">https://discere.tech/java-memory-model-concurrency-cheat-sheet</guid><dc:creator><![CDATA[Tarun Bansal]]></dc:creator><pubDate>Fri, 28 Aug 2026 20:31:12 GMT</pubDate><content:encoded><![CDATA[<p>A practical reference covering JVM memory internals and Java concurrency — heap structure, the Java Memory Model, thread lifecycle, synchronization primitives, executors, <code>CompletableFuture</code>, and the concurrency bugs that bite hardest.</p>
<hr />
<h2>Memory</h2>
<h3>JVM Heap Generations</h3>
<p>The JVM heap is divided into generations. Young Gen holds newly allocated objects and is collected frequently. Surviving objects are promoted to Old Gen (Tenured). Metaspace (Java 8+) replaced PermGen and holds class metadata — it grows dynamically and is not part of the heap.</p>
<pre><code class="language-plaintext">┌─────────────────────────────────────────────────────────┐
│                        JVM HEAP                         │
│  ┌──────────────────────────────┐  ┌──────────────────┐ │
│  │         Young Gen            │  │    Old Gen       │ │
│  │  ┌───────┐ ┌─────┐ ┌─────┐  │  │   (Tenured)      │ │
│  │  │ Eden  │ │ S0  │ │ S1  │  │──▶│                  │ │
│  │  │       │ │     │ │     │  │  │  Long-lived       │ │
│  │  └───────┘ └─────┘ └─────┘  │  │  objects         │ │
│  └──────────────────────────────┘  └──────────────────┘ │
└─────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│              Metaspace (off-heap)                        │
│         Class metadata, static fields                    │
└─────────────────────────────────────────────────────────┘
</code></pre>
<h3>Stack vs Heap</h3>
<p>Each thread has its own stack holding stack frames (local variables, method calls). Objects always live on the heap. Stack overflow = infinite recursion or a call stack that's too deep.</p>
<pre><code class="language-java">// Stack: holds primitives and references (not objects)
public void calculate() {
    int x = 5;           // x lives on THIS thread's stack
    String s = "hello";  // reference 's' on stack, String object on heap
    Object obj = new Object(); // obj ref on stack, Object on heap
}

// Stack overflow example
public int infinite(int n) {
    return infinite(n + 1); // StackOverflowError — no base case
}
</code></pre>
<h3>Java Memory Model (JMM) — Happens-Before Rules</h3>
<p>The JMM defines happens-before (HB) rules that guarantee memory visibility between threads:</p>
<ul>
<li><p><strong>Monitor unlock HB lock</strong> — releasing a lock flushes writes; acquiring reads fresh values</p>
</li>
<li><p><strong>Volatile write HB read</strong> — a volatile write makes all prior writes visible to any subsequent reader</p>
</li>
<li><p><strong>Thread.start() HB first action</strong> — parent thread's state is visible to the new thread</p>
</li>
<li><p><strong>Thread.join() HB caller</strong> — child thread's writes are visible after <code>join()</code></p>
</li>
</ul>
<pre><code class="language-java">// Rule 1: Monitor unlock HB lock
int x = 0;
synchronized (lock) { x = 42; }   // UNLOCK
// --- another thread ---
synchronized (lock) {              // LOCK
    System.out.println(x);         // guaranteed to see 42
}

// Rule 2: Volatile write HB read
volatile boolean ready = false;
int data = 0;
// Thread A
data = 100;
ready = true;   // volatile WRITE
// Thread B
if (ready) use(data);   // sees data = 100

// Rule 3: Thread start HB first action
config = 99;
Thread t = new Thread(() -&gt; use(config)); // sees 99
t.start();
</code></pre>
<h3>Off-Heap Memory</h3>
<p>Off-heap memory lives outside the JVM heap — never touched by GC. Used by Kafka clients, Netty, and high-throughput banking apps. Must be explicitly freed. Two APIs: <code>DirectByteBuffer</code> (standard) and <code>sun.misc.Unsafe</code> (raw pointer arithmetic).</p>
<pre><code class="language-java">// DirectByteBuffer — standard API
ByteBuffer buf = ByteBuffer.allocateDirect(1024 * 1024); // 1 MB off-heap
buf.putInt(0, 42);
int val = buf.getInt(0);
// Released when ByteBuffer wrapper is GC'd (via Cleaner) — not immediate!

// Force immediate free (pre-Java 9)
((sun.nio.ch.DirectBuffer) buf).cleaner().clean();

// Unsafe — raw pointer arithmetic, no bounds checking
Unsafe unsafe = getUnsafe();
long address = unsafe.allocateMemory(1024);
unsafe.putLong(address, 0xDEADBEEFL);
long result = unsafe.getLong(address);
try {
    riskyOperation();
} finally {
    unsafe.freeMemory(address);  // MUST free — GC will never do this
}

// Common leak: exhausting direct memory
for (int i = 0; i &lt; 100_000; i++) {
    ByteBuffer.allocateDirect(1024 * 1024); // OOM — GC too slow
}
</code></pre>
<h3>Common Memory Leak Patterns</h3>
<pre><code class="language-java">// 1. Static collections holding references
static List&lt;byte[]&gt; cache = new ArrayList&lt;&gt;();
cache.add(new byte[1024 * 1024]); // never removed → lives forever

// 2. Unclosed streams/connections
InputStream in = new FileInputStream("data.txt");
// forgot in.close() → file descriptor + buffer leaked

// 3. ThreadLocal not removed in thread pools
static ThreadLocal&lt;UserContext&gt; ctx = new ThreadLocal&lt;&gt;();
ctx.set(new UserContext()); // thread is reused from pool
// ctx.remove() never called → old context survives next request

// 4. Listeners not deregistered
eventBus.register(this);  // adds reference to eventBus's list
// object can't be GC'd even after you're "done" with it
// Fix: always call eventBus.unregister(this) in cleanup

// 5. Non-static inner class holding outer reference
class Outer {
    byte[] hugeData = new byte[10_000_000];
    class Inner implements Runnable {  // holds implicit ref to Outer
        public void run() { /* ... */ }
    }
}
executor.submit(new Outer().new Inner());
// Outer + hugeData pinned until task completes
</code></pre>
]]></content:encoded></item><item><title><![CDATA[Kafka Core Concepts: From Consumer Groups to High Throughput]]></title><description><![CDATA[A practical guide to consumer groups, offsets, Schema Registry, producer internals and the design choices that make Kafka fast.

Kafka Core Concepts: From Consumer Groups to High Throughput
Kafka is o]]></description><link>https://discere.tech/kafka-core-concepts-from-consumer-groups-to-high-throughput</link><guid isPermaLink="true">https://discere.tech/kafka-core-concepts-from-consumer-groups-to-high-throughput</guid><dc:creator><![CDATA[Tarun Bansal]]></dc:creator><pubDate>Fri, 28 Aug 2026 19:25:12 GMT</pubDate><content:encoded><![CDATA[<hr />
<p>A practical guide to consumer groups, offsets, Schema Registry, producer internals and the design choices that make Kafka fast.</p>
<hr />
<h1>Kafka Core Concepts: From Consumer Groups to High Throughput</h1>
<p>Kafka is often introduced as a distributed event-streaming platform, but that description does not explain how it behaves under load—or why seemingly small configuration changes can alter its delivery guarantees.</p>
<p>To use Kafka confidently, you need a mental model that connects five areas:</p>
<ol>
<li><p>How consumer groups divide work</p>
</li>
<li><p>How offsets track progress</p>
</li>
<li><p>How schemas evolve without breaking applications</p>
</li>
<li><p>What happens inside the producer after <code>send()</code></p>
</li>
<li><p>Why a disk-backed platform can still achieve high throughput</p>
</li>
</ol>
<p>This article connects those ideas and highlights the production details that simplified explanations often miss.</p>
<blockquote>
<p><strong>Version note:</strong> Configuration defaults in this article are based on Apache Kafka 4.x. In particular, <code>linger.ms</code> defaults to <code>5</code> ms from Kafka 4.0 onward; older references commonly show <code>0</code>.</p>
</blockquote>
<h2>1. Consumer Groups and Partition Assignment</h2>
<p>A <strong>consumer group</strong> is a set of consumers that cooperate to read one or more topics. Within a group, every subscribed partition has exactly one active consumer at a time.</p>
<p>Consider an <code>orders</code> topic with four partitions:</p>
<pre><code class="language-text">Topic: orders

P0 ──► Consumer A
P1 ──► Consumer B       Consumer group: order-svc
P2 ──► Consumer C
P3 ──► Consumer D
</code></pre>
<p>If another consumer joins the same group, Kafka rebalances the assignments. With four partitions and five consumers, one consumer remains idle because Kafka cannot assign the same partition to two active consumers in one group.</p>
<p>The scaling rule is therefore straightforward:</p>
<blockquote>
<p>Useful consumer parallelism within a group is bounded by the number of assigned partitions.</p>
</blockquote>
<p>Adding partitions can increase potential parallelism, but it is not a free operation. More partitions mean more metadata, files, replication work and rebalance complexity. They can also change key distribution for newly produced records.</p>
<h3>Partition ownership does not eliminate duplicates</h3>
<p>The one-consumer-per-partition rule prevents two current members of the same group from intentionally processing the partition simultaneously. It does <strong>not</strong> guarantee exactly-once business processing.</p>
<p>Suppose a consumer:</p>
<ol>
<li><p>Reads record 42</p>
</li>
<li><p>Updates a database</p>
</li>
<li><p>Crashes before committing its new offset</p>
</li>
</ol>
<p>After reassignment, another consumer resumes from the previous committed position and processes record 42 again. The application must therefore use idempotent processing, deduplication or an appropriate transactional pattern.</p>
<h3>Java consumer example</h3>
<pre><code class="language-java">Properties props = new Properties();
props.put(ConsumerConfig.BOOTSTRAP_SERVERS_CONFIG, "broker1:9092");
props.put(ConsumerConfig.GROUP_ID_CONFIG, "order-svc");
props.put(
    ConsumerConfig.KEY_DESERIALIZER_CLASS_CONFIG,
    StringDeserializer.class
);
props.put(
    ConsumerConfig.VALUE_DESERIALIZER_CLASS_CONFIG,
    StringDeserializer.class
);
props.put(ConsumerConfig.AUTO_OFFSET_RESET_CONFIG, "earliest");

try (KafkaConsumer&lt;String, String&gt; consumer = new KafkaConsumer&lt;&gt;(props)) {
    consumer.subscribe(List.of("orders"));

    while (true) {
        ConsumerRecords&lt;String, String&gt; records =
            consumer.poll(Duration.ofMillis(100));

        for (ConsumerRecord&lt;String, String&gt; record : records) {
            process(record);
        }
    }
}
</code></pre>
<h2>2. Offsets and Commit Management</h2>
<p>An <strong>offset</strong> is a monotonically increasing position within a partition. It is unique only within that partition.</p>
<p>For example:</p>
<pre><code class="language-text">Partition 0

Record offset:     0   1   2   3   4   5   6
                              ▲
                         processed 4

Committed position: 5
Restart position:   5
</code></pre>
<p>Kafka commits the offset of the <strong>next record to consume</strong>. If offset <code>4</code> has been processed successfully, the committed value should normally be <code>5</code>.</p>
<p>For consumer groups, committed positions are stored in Kafka’s internal <code>__consumer_offsets</code> topic.</p>
<h3>Position and committed position are different</h3>
<p>The consumer’s current position advances as <code>poll()</code> returns data. The committed position changes only when an offset commit succeeds.</p>
<p>That distinction creates three common processing models:</p>
<table>
<thead>
<tr>
<th>Commit timing</th>
<th>Possible failure outcome</th>
<th>Typical semantic</th>
</tr>
</thead>
<tbody><tr>
<td>Before processing</td>
<td>Work can be skipped after a crash</td>
<td>At-most-once</td>
</tr>
<tr>
<td>After processing</td>
<td>Completed work can be repeated</td>
<td>At-least-once</td>
</tr>
<tr>
<td>Kafka transaction covering output and offsets</td>
<td>Kafka output and consumed positions commit atomically</td>
<td>Exactly-once within the supported Kafka scope</td>
</tr>
</tbody></table>
<h3>Manual offset commit</h3>
<pre><code class="language-java">props.put(ConsumerConfig.ENABLE_AUTO_COMMIT_CONFIG, "false");

ConsumerRecords&lt;String, String&gt; records =
    consumer.poll(Duration.ofMillis(100));

Map&lt;TopicPartition, OffsetAndMetadata&gt; offsets = new HashMap&lt;&gt;();

for (ConsumerRecord&lt;String, String&gt; record : records) {
    process(record);

    offsets.put(
        new TopicPartition(record.topic(), record.partition()),
        new OffsetAndMetadata(record.offset() + 1)
    );
}

consumer.commitSync(offsets);
</code></pre>
<p>In real applications, be careful when processing records concurrently. Committing the largest offset while an earlier record is still running can cause Kafka to skip that unfinished work after a restart.</p>
<h2>3. Schema Registry and Safe Schema Evolution</h2>
<p>A <strong>Schema Registry</strong> stores and versions event schemas such as Avro, Protobuf and JSON Schema. It creates an explicit contract between producers and consumers.</p>
<p>The typical flow is:</p>
<pre><code class="language-text">Producer ──register schema──► Schema Registry
Producer ◄────schema ID────── Schema Registry

Producer ──serialized record──► Kafka

Consumer ──resolve schema ID──► Schema Registry
Consumer ◄────writer schema──── Schema Registry
</code></pre>
<p>Schema lookups are cached by serializers and deserializers. A consumer does not normally make a registry request for every record.</p>
<h3>Confluent wire format</h3>
<p>The traditional Confluent payload-prefix format is:</p>
<pre><code class="language-text">[ magic byte: 0x00 ][ schema ID: 4 bytes ][ encoded payload ]
</code></pre>
<p>Current Confluent versions can also carry the schema identifier in a Kafka record header. The payload-prefix representation remains the default for the standard serializers.</p>
<h3>Compatibility modes</h3>
<table>
<thead>
<tr>
<th>Mode</th>
<th>Guarantee</th>
<th>Usual deployment direction</th>
</tr>
</thead>
<tbody><tr>
<td><code>BACKWARD</code></td>
<td>New readers can read data written with the previous schema</td>
<td>Consumers first</td>
</tr>
<tr>
<td><code>FORWARD</code></td>
<td>Previous readers can read data written with the new schema</td>
<td>Producers first</td>
</tr>
<tr>
<td><code>FULL</code></td>
<td>Both backward and forward compatible</td>
<td>Either direction, subject to checks</td>
</tr>
<tr>
<td><code>NONE</code></td>
<td>No compatibility validation</td>
<td>Avoid in production</td>
</tr>
</tbody></table>
<p>Compatibility is evaluated according to the schema format. Avro, Protobuf and JSON Schema do not have identical evolution rules.</p>
<h3>Avro evolution example</h3>
<p>Adding a field with a default allows a new Avro reader to read older data that does not contain that field:</p>
<pre><code class="language-json">{
  "type": "record",
  "name": "OrderPlaced",
  "namespace": "com.investment.orders.v1",
  "fields": [
    { "name": "orderId", "type": "string" },
    { "name": "portfolioId", "type": "string" },
    { "name": "amount", "type": "double" },
    { "name": "currency", "type": "string", "default": "EUR" },
    { "name": "email", "type": ["null", "string"], "default": null }
  ]
}
</code></pre>
<p>The <code>email</code> field is genuinely nullable because its type is a union containing <code>null</code>, and its default matches the first union member.</p>
<p>A textual field rename is normally breaking. Avro aliases can support controlled migrations, but they should be tested against the actual compatibility mode, serializer and consumer implementations rather than treated as universally safe.</p>
<h3>Producer configuration</h3>
<pre><code class="language-java">props.put(
    ProducerConfig.VALUE_SERIALIZER_CLASS_CONFIG,
    KafkaAvroSerializer.class
);
props.put("schema.registry.url", "http://schema-registry:8081");
</code></pre>
<h3>Compatibility check</h3>
<pre><code class="language-bash">curl -X POST \
  http://schema-registry:8081/compatibility/subjects/orders.placed-value/versions/latest \
  -H 'Content-Type: application/vnd.schemaregistry.v1+json' \
  -d '{"schema":"{...escaped candidate schema...}"}'
</code></pre>
<p>Set compatibility for a subject:</p>
<pre><code class="language-bash">curl -X PUT \
  http://schema-registry:8081/config/orders.placed-value \
  -H 'Content-Type: application/vnd.schemaregistry.v1+json' \
  -d '{"compatibility":"FULL"}'
</code></pre>
<p>Run this compatibility check in CI before deploying a producer or consumer that introduces a new schema version.</p>
<h2>4. What Happens Inside the Kafka Producer?</h2>
<p><code>KafkaProducer.send()</code> starts a pipeline rather than performing a synchronous network write:</p>
<pre><code class="language-text">send()
  │
  ▼
Serialize key and value
  │
  ▼
Choose partition
  │
  ▼
RecordAccumulator: per-partition batches
  │
  ▼
Background Sender and NetworkClient
  │
  ▼
Kafka broker
</code></pre>
<p>The main stages are:</p>
<ol>
<li><p><strong>Serialization:</strong> Key and value serializers convert application objects to bytes.</p>
</li>
<li><p><strong>Partition selection:</strong> An explicit partition takes precedence. Otherwise, the producer applies key-based or sticky partitioning behaviour.</p>
</li>
<li><p><strong>Accumulation:</strong> Records enter batches associated with their target partitions.</p>
</li>
<li><p><strong>Drain:</strong> The background sender creates requests when batches become ready.</p>
</li>
<li><p><strong>Completion:</strong> Broker responses complete futures and invoke callbacks.</p>
</li>
</ol>
<h3>Important producer settings</h3>
<table>
<thead>
<tr>
<th>Setting</th>
<th>Kafka 4.x default</th>
<th>Purpose</th>
</tr>
</thead>
<tbody><tr>
<td><code>batch.size</code></td>
<td>16,384 bytes</td>
<td>Upper bound for a normal record batch per partition</td>
</tr>
<tr>
<td><code>linger.ms</code></td>
<td>5 ms</td>
<td>Maximum batching delay under normal conditions</td>
</tr>
<tr>
<td><code>buffer.memory</code></td>
<td>33,554,432 bytes</td>
<td>Approximate total producer buffer budget</td>
</tr>
<tr>
<td><code>compression.type</code></td>
<td><code>none</code></td>
<td>Compression applied to complete batches</td>
</tr>
<tr>
<td><code>max.in.flight.requests.per.connection</code></td>
<td>5</td>
<td>Unacknowledged requests allowed per broker connection</td>
</tr>
</tbody></table>
<p><code>send()</code> is asynchronous, but it is not guaranteed to return immediately. It can block while waiting for metadata or buffer space, up to <code>max.block.ms</code>. Serialization also runs on the calling thread.</p>
<h2>5. Tuning for Throughput and Latency</h2>
<p>There is no universal “fastest” configuration. Measure end-to-end p95 and p99 latency, throughput, batch-size average, compression rate, retries, buffer exhaustion and broker request latency under a realistic workload.</p>
<h3>Durable high-throughput baseline</h3>
<pre><code class="language-java">Properties props = new Properties();

props.put(ProducerConfig.BATCH_SIZE_CONFIG, 65_536);
props.put(ProducerConfig.LINGER_MS_CONFIG, 10);
props.put(ProducerConfig.COMPRESSION_TYPE_CONFIG, "lz4");
props.put(ProducerConfig.BUFFER_MEMORY_CONFIG, 67_108_864L);
props.put(ProducerConfig.ACKS_CONFIG, "all");
props.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true);
props.put(
    ProducerConfig.MAX_IN_FLIGHT_REQUESTS_PER_CONNECTION,
    5
);
</code></pre>
<p>These values are starting points, not prescriptions. Benchmark <code>lz4</code> and <code>zstd</code> with your payload and CPU budget.</p>
<h3>Latency-sensitive baseline</h3>
<pre><code class="language-java">props.put(ProducerConfig.LINGER_MS_CONFIG, 0);
props.put(ProducerConfig.BATCH_SIZE_CONFIG, 16_384);
props.put(ProducerConfig.COMPRESSION_TYPE_CONFIG, "none");
props.put(ProducerConfig.ACKS_CONFIG, "all");
props.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true);
</code></pre>
<p>For financially important events, changing to <code>acks=1</code> merely to reduce latency is usually an unsafe default. Benchmark <code>acks=all</code> with idempotence first. Any weaker durability guarantee should be an explicit business-risk decision.</p>
<h3>Safe asynchronous callback</h3>
<pre><code class="language-java">producer.send(record, (metadata, exception) -&gt; {
    if (exception != null) {
        // metadata may be null when the send fails
        log.error("Failed to publish record", exception);
        alertOrRouteForRecovery(record, exception);
        return;
    }

    log.debug(
        "Published to {}-{} at offset {}",
        metadata.topic(),
        metadata.partition(),
        metadata.offset()
    );
});
</code></pre>
<p>Keep callbacks lightweight because they execute on the producer’s I/O thread.</p>
<h2>6. Why Is Kafka Fast Even Though It Uses Disk?</h2>
<p>Kafka’s throughput comes from several mechanisms working together.</p>
<h3>Sequential log access</h3>
<p>Kafka writes partition logs as append-oriented files. Sequential access minimizes random seeks and works well with filesystem prefetching and modern storage devices.</p>
<p>This is efficient, but “disk is as fast as RAM” is too absolute. Real performance depends on the storage medium, filesystem, page-cache hit rate, flush policy and workload.</p>
<h3>Operating-system page cache</h3>
<p>Kafka relies heavily on the operating system’s page cache rather than maintaining a separate application-level cache of log data. Recently written or read pages can be served from memory, while the OS manages eviction and writeback.</p>
<h3>Zero-copy transfer</h3>
<p>Where the operating system and connection path allow it, Kafka can use <code>sendfile</code>-style transfer to avoid copying record bytes through the JVM heap.</p>
<pre><code class="language-text">Conventional user-space path:

Storage → kernel page cache → application buffer → socket → NIC

Zero-copy path:

Storage → kernel page cache ──────────────────────────→ NIC
                         transfer stays in kernel path
</code></pre>
<p>The exact number of physical copies depends on the operating system, TLS configuration and hardware capabilities. Zero-copy is best understood as reducing application-level copying and CPU overhead.</p>
<h3>Batching</h3>
<p>Producers send multiple records in a single batch, and consumers fetch records in batches. This amortizes network round trips, checksums, compression and protocol overhead across many records.</p>
<h3>Partition-level parallelism</h3>
<p>Partitions distribute storage, replication, production and consumption across brokers and consumers. Effective parallelism is still limited by key distribution: a hot key can create a hot partition even when the topic has many partitions.</p>
<h2>Final Mental Model</h2>
<p>The major Kafka concepts fit together as one continuous lifecycle:</p>
<pre><code class="language-text">Producer
  └─ serializes, partitions and batches records
         │
         ▼
Kafka partition log
  └─ stores ordered records efficiently
         │
         ▼
Consumer group
  └─ assigns each partition to one active member
         │
         ▼
Offset commit
  └─ records the next restart position

Schema Registry governs the contract across the full flow.
</code></pre>
<p>The most useful production lessons are these:</p>
<ul>
<li><p>Partition assignment enables parallelism but does not eliminate duplicate processing.</p>
</li>
<li><p>An offset is a restart position, not proof that every external side effect succeeded.</p>
</li>
<li><p>Compatibility checks should be automated before deployment.</p>
</li>
<li><p>Producer tuning always trades memory, batching, latency and durability.</p>
</li>
<li><p>Kafka is fast because sequential access, page cache, efficient transfer and batching reinforce one another.</p>
</li>
</ul>
<h2>References</h2>
<ul>
<li><p><a href="https://kafka.apache.org/41/configuration/producer-configs/">Apache Kafka 4.1 Producer Configurations</a></p>
</li>
<li><p><a href="https://kafka.apache.org/41/design/design/">Apache Kafka 4.1 Design</a></p>
</li>
<li><p><a href="https://kafka.apache.org/41/javadoc/org/apache/kafka/clients/producer/KafkaProducer.html">Apache Kafka <code>KafkaProducer</code> API</a></p>
</li>
<li><p><a href="https://docs.confluent.io/platform/current/schema-registry/fundamentals/serdes-develop/index.html">Confluent Schema Registry SerDes</a></p>
</li>
<li><p><a href="https://docs.confluent.io/platform/current/schema-registry/fundamentals/serdes-develop/serdes-avro.html">Confluent Avro Serializer and Deserializer</a></p>
</li>
</ul>
<hr />
<p><em>If you found this useful, follow</em> <em><strong>Systems in Practice</strong></em> <em>for more articles on Kafka, Java, distributed systems and reliable backend architecture.</em></p>
]]></content:encoded></item><item><title><![CDATA[Designing Real-Time Pricing, Quotes and Orders on AWS]]></title><description><![CDATA[Real-time pricing in an investment platform looks simple from the customer’s perspective: open a portfolio, view its current value and submit an order.
Behind that apparently simple flow is a difficul]]></description><link>https://discere.tech/designing-real-time-pricing-quotes-and-orders-on-aws</link><guid isPermaLink="true">https://discere.tech/designing-real-time-pricing-quotes-and-orders-on-aws</guid><dc:creator><![CDATA[Tarun Bansal]]></dc:creator><pubDate>Fri, 28 Aug 2026 15:44:59 GMT</pubDate><content:encoded><![CDATA[<p>Real-time pricing in an investment platform looks simple from the customer’s perspective: open a portfolio, view its current value and submit an order.</p>
<p>Behind that apparently simple flow is a difficult distributed-systems problem. Market prices change continuously, messages can arrive out of order, data vendors can fail, and a price shown to a customer may become stale within seconds.</p>
<p>The central design question is:</p>
<blockquote>
<p>How can we process continuously changing market prices while guaranteeing that an order uses the exact price accepted by the customer?</p>
</blockquote>
<p>The answer begins by separating two fundamentally different types of data.</p>
<h2>Two Data Paths with Different Responsibilities</h2>
<p>The platform should maintain two independent data paths:</p>
<ol>
<li><p><strong>Market-price data</strong> — high-volume, continuously changing and optimized for fast reads and writes.</p>
</li>
<li><p><strong>Quote and order data</strong> — lower-volume, transactional, immutable and auditable.</p>
</li>
</ol>
<p>The latest market price and an accepted customer price are not the same business fact.</p>
<p>The pricing platform tells us what an instrument is worth now. The quote and ordering platform records what price the business offered and what the customer accepted.</p>
<p>Trying to store both in the same database creates unnecessary coupling. A sudden increase in market-data traffic could then affect order processing, even though orders have much stronger consistency requirements.</p>
<h2>AWS-Native Architecture</h2>
<p>The architecture can be built using the following components:</p>
<ul>
<li><p>Vendor adapters running on Amazon EKS or ECS</p>
</li>
<li><p>Amazon MSK for durable price-event streaming</p>
</li>
<li><p>DynamoDB for the latest price of each instrument</p>
</li>
<li><p>Amazon Aurora PostgreSQL for quotes and orders</p>
</li>
<li><p>Amazon S3 for historical price events and replay</p>
</li>
<li><p>A transactional outbox for reliable business-event publication</p>
</li>
<li><p>CloudWatch, OpenTelemetry and Prometheus for observability</p>
</li>
</ul>
<p>The logical flow is:</p>
<pre><code class="language-text">Primary and fallback vendors
        ↓
Vendor adapters
        ↓
Amazon MSK
        ↓
Price processor
        ↓
DynamoDB latest-price store
        ↓
Portfolio and Quote services
        ↓
Aurora quote and order database
</code></pre>
<p>Amazon MSK decouples the external vendors from the internal pricing platform. It absorbs bursts, supports replay and allows price-processing consumers to scale independently.</p>
<p>Events should be partitioned by <code>instrumentId</code>. This preserves ordering for a particular instrument without forcing the entire market-data stream through one partition. However, partitioning by instrument can create hot partitions for high-volume instruments such as AAPL, TSLA or index futures. The deployment should therefore provision enough MSK partitions and consumers for the expected concentration, monitor per-partition throughput and lag, and define an explicit policy for skew. If a single instrument can exceed the capacity of one partition, the design must either allocate that instrument to a dedicated feed or shard its events while preserving ordering through a downstream per-instrument sequencer. The default design accepts one ordered partition per instrument when its throughput remains within the partition limit.</p>
<h2>Why DynamoDB for Current Prices?</h2>
<p>The latest-price store has a simple but demanding workload:</p>
<ul>
<li><p>frequent updates;</p>
</li>
<li><p>access by instrument identifier;</p>
</li>
<li><p>predictable low-latency reads;</p>
</li>
<li><p>independent scaling;</p>
</li>
<li><p>no requirement for relational joins.</p>
</li>
</ul>
<p>A representative price record might contain:</p>
<pre><code class="language-json">{
  "instrumentId": "AAPL",
  "bid": 229.12,
  "ask": 229.18,
  "currency": "USD",
  "vendor": "PRIMARY",
  "vendorTimestamp": "2026-08-28T15:30:12.341Z",
  "receivedAt": "2026-08-28T15:30:12.366Z",
  "sourceSequence": 91827364,
  "priceVersion": "PRIMARY:91827364",
  "status": "LIVE"
}
</code></pre>
<p>DynamoDB is well suited to this access pattern because the processor can continuously replace the latest state for each instrument.</p>
<p>Historical prices do not need to remain in the same table indefinitely. The original events can be retained in Kafka for a configured period and archived to Amazon S3 for audit, analytics and recovery.</p>
<h2>Handling Duplicate and Out-of-Order Prices</h2>
<p>Distributed messaging systems normally provide at-least-once delivery. A price processor must therefore expect both duplicate and out-of-order events.</p>
<p>A vendor-provided sequence number is the best ordering mechanism. If one is unavailable, the system may use a vendor timestamp combined with a deterministic tie-breaker.</p>
<p>A timestamp alone is insufficient because:</p>
<ul>
<li><p>vendor clocks may not be synchronized;</p>
</li>
<li><p>two events may have identical timestamps;</p>
</li>
<li><p>timestamps may use different precision;</p>
</li>
<li><p>network delays can change arrival order;</p>
</li>
<li><p>different vendors may have different notions of event time.</p>
</li>
</ul>
<p>The price processor should use a conditional update so that an older event cannot overwrite a newer price.</p>
<p>Ordering should generally be evaluated within the combination of:</p>
<pre><code class="language-text">instrument + vendor + feed
</code></pre>
<p>Sequence numbers from two different vendors should not be directly compared unless their contract explicitly guarantees a common sequence.</p>
<p>The conditional update protects the latest-price record from out-of-order writes; it does not guarantee that a later reader will observe its own write. The Quote Service must therefore use the appropriate DynamoDB read consistency for the freshness-sensitive path. For a single-item read of the hot latest-price key, a strongly consistent read should be used when the service must immediately observe the latest accepted record. This is separate from the ordering guarantee: conditional writes prevent stale events from winning, while read consistency determines what a subsequent Quote Service read can observe.</p>
<h2>Creating a Customer Quote</h2>
<p>When a customer asks for a portfolio price, the Quote Service reads the required instrument prices from DynamoDB. These reads are observations taken at particular instants; they do not form an atomic snapshot across all instruments or across DynamoDB and Aurora. A price may change after it is read and before the quote is written. That is acceptable and intentional.</p>
<p>The safety of the quote does not come from promising that the price remained unchanged during the operation. It comes from storing the exact price and its <code>priceVersion</code> in the immutable Aurora quote. The <code>priceVersion</code> identifies the source sequence or equivalent version that was captured and used for calculation. The resulting quote therefore records which market observation became the business commitment, even if the latest-price record changes before or immediately after the Aurora transaction.</p>
<p>Before creating a quote, it validates:</p>
<ul>
<li><p>every required instrument has a price;</p>
</li>
<li><p>each price is sufficiently fresh;</p>
</li>
<li><p>currencies and market sessions are correct;</p>
</li>
<li><p>no instrument has been suspended;</p>
</li>
<li><p>all pricing rules can be applied.</p>
</li>
</ul>
<p>The freshness check should use a platform-controlled time basis. <code>receivedAt</code>, compared with the Quote Service’s trusted server time, should normally determine whether the record is within the two-second operational freshness window. <code>vendorTimestamp</code> should be retained for audit and event-time analysis, but should not normally be used as the sole basis for a strict freshness SLA because vendor clock skew can make a price appear newer or older than it is. If vendor time is used for any business rule, clock-skew bounds and validation must be explicitly defined.</p>
<p>The service should evaluate freshness for every instrument in the portfolio. Under the default all-or-nothing policy, a portfolio is unquotable if any required instrument is missing, suspended, invalid or outside the freshness threshold. There is no degraded or partial quote unless a separate product policy explicitly defines one. This policy should be confirmed with the product and risk owners because a portfolio containing many instruments has a higher probability that at least one instrument will be stale. If partial quoting is later desired, it must define how omitted or stale instruments affect totals, customer acceptance and order creation; it must not be introduced implicitly.</p>
<p>It then stores an immutable quote in Aurora PostgreSQL.</p>
<p>The quote should include:</p>
<ul>
<li><p>quote ID;</p>
</li>
<li><p>customer and portfolio IDs;</p>
</li>
<li><p>instrument identifiers and quantities;</p>
</li>
<li><p>exact instrument prices;</p>
</li>
<li><p>price versions;</p>
</li>
<li><p>exchange rates;</p>
</li>
<li><p>fees;</p>
</li>
<li><p>calculated total;</p>
</li>
<li><p>pricing-rule version;</p>
</li>
<li><p>quote creation time;</p>
</li>
<li><p>expiry time;</p>
</li>
<li><p>source metadata.</p>
</li>
</ul>
<p>Storing only the final portfolio total would not be sufficient. The platform must be able to explain exactly how that amount was calculated.</p>
<h2>Price Freshness and Quote Expiry Are Different</h2>
<p>Two related but separate time limits are involved.</p>
<p><strong>Price freshness</strong> determines whether a market price is recent enough to create a new quote.</p>
<p><strong>Quote validity</strong> determines how long the business promises to honour a quote after creating it.</p>
<p>For example:</p>
<pre><code class="language-text">Maximum acceptable market-price age: 2 seconds
Customer quote validity: 10 seconds
</code></pre>
<p>A three-second-old market price cannot be used to create a new quote, even if a successfully created quote remains valid for ten seconds.</p>
<p>This distinction is important because the first rule protects pricing quality, while the second represents a commercial commitment to the customer.</p>
<p>The freshness calculation should use <code>receivedAt</code> and a trusted server-side clock, with an explicitly documented allowance for infrastructure clock skew. <code>vendorTimestamp</code> remains useful for diagnosing vendor latency and validating feed behaviour, but vendor-supplied time must not silently become the authoritative clock for the freshness SLA.</p>
<h2>Accepting a Quote</h2>
<p>The ordering domain should own quote acceptance.</p>
<p>When the customer accepts a quote, the request includes the quote ID and an idempotency key. The Order Service then performs a single Aurora transaction:</p>
<ol>
<li><p>Verify that the quote belongs to the customer.</p>
</li>
<li><p>Check that its status is <code>ACTIVE</code>.</p>
</li>
<li><p>Check that it has not expired.</p>
</li>
<li><p>Mark it as <code>ACCEPTED</code>.</p>
</li>
<li><p>Create the order using the stored quote prices.</p>
</li>
<li><p>Insert an event into the transactional outbox.</p>
</li>
<li><p>Commit the transaction.</p>
</li>
</ol>
<p>A conditional update can enforce the decision atomically:</p>
<pre><code class="language-sql">UPDATE quote
SET status = 'ACCEPTED',
    accepted_at = CURRENT_TIMESTAMP
WHERE quote_id = :quoteId
  AND status = 'ACTIVE'
  AND expires_at &gt;= CURRENT_TIMESTAMP;
</code></pre>
<p>If one row is updated, acceptance succeeded. If no rows are updated, the quote was expired, cancelled or previously accepted.</p>
<p>The server-side database clock must determine expiry. The browser countdown is only a user-interface aid and cannot be authoritative.</p>
<h2>Why Not Use a Distributed Transaction?</h2>
<p>Creating a quote involves reading from DynamoDB and writing to Aurora. There is no single ACID transaction across these two databases.</p>
<p>That is acceptable because the market price is an observation, while the stored quote is a business commitment.</p>
<p>The read and write are therefore intentionally non-atomic. Between reading a price from DynamoDB and writing the quote to Aurora, the latest price may update. The system does not attempt to prevent that update or claim that the read remained current. Instead, the quote stores the exact values and <code>priceVersion</code> values captured by the Quote Service. Those versions make the quote auditable and safe: they identify the market observations used for the commitment.</p>
<p>The DynamoDB conditional update protects the latest-price record from an older event overwriting a newer event. It does not provide read-your-write consistency to the Quote Service. That guarantee must come from the selected DynamoDB read mode, with strongly consistent reads used for the freshness-sensitive latest-price lookup where required.</p>
<p>The database transaction is required only when the quote becomes an order. Quote acceptance, order creation and outbox insertion all belong to the same consistency boundary and therefore occur in Aurora.</p>
<h2>Reliable Business-Event Publication</h2>
<p>The transactional outbox should use the standard pattern:</p>
<ol>
<li><p>The Aurora transaction changes the quote and order state.</p>
</li>
<li><p>The same transaction inserts an outbox row containing the business event, event ID, aggregate ID, payload and publication status.</p>
</li>
<li><p>A separate publisher polls the outbox or consumes database change data capture.</p>
</li>
<li><p>The publisher sends the event to the target broker or service.</p>
</li>
<li><p>Successful publication is recorded, and retries continue for failures.</p>
</li>
</ol>
<p>The publisher must be idempotent because a crash can occur after publication but before the outbox row is marked as published. Consumers should deduplicate using the event ID.</p>
<p>This pattern makes both of the following recoverable:</p>
<ul>
<li><p>if Aurora commits but the response is lost, a retry can return the already-created order using the idempotency key;</p>
</li>
<li><p>if event publication fails after the Aurora commit, the durable outbox row remains available for later retry.</p>
</li>
</ul>
<p>The outbox is therefore part of the Aurora consistency boundary, not merely an asynchronous best-effort notification mechanism.</p>
<h2>Vendor Failover</h2>
<p>The second price vendor should not simply overwrite the primary vendor whenever it provides a later timestamp.</p>
<p>Failover should be controlled through explicit states:</p>
<pre><code class="language-text">PRIMARY_ACTIVE
FALLBACK_PENDING
FALLBACK_ACTIVE
PRIMARY_PENDING
</code></pre>
<p>Before activating the fallback feed, the platform should validate:</p>
<ul>
<li><p>feed heartbeat;</p>
</li>
<li><p>instrument coverage;</p>
</li>
<li><p>price freshness;</p>
</li>
<li><p>bid/ask spread;</p>
</li>
<li><p>deviation from the last trusted primary price;</p>
</li>
<li><p>timestamp or sequence progression;</p>
</li>
<li><p>currency and trading-session information.</p>
</li>
</ul>
<p>The same validation is required before switching back to the primary vendor. A recently recovered feed should remain stable for a defined period before it becomes authoritative again.</p>
<h2>Important Failure Scenarios</h2>
<p>A robust design must define its behaviour before these failures occur:</p>
<table>
<thead>
<tr>
<th>Scenario</th>
<th>Expected behaviour</th>
</tr>
</thead>
<tbody><tr>
<td>Price changes after quote creation</td>
<td>Honour the stored quote until it expires</td>
</tr>
<tr>
<td>Price changes between DynamoDB read and Aurora write</td>
<td>Store the captured values and <code>priceVersion</code>; do not require the price to remain unchanged</td>
</tr>
<tr>
<td>DynamoDB read observes an older record</td>
<td>Use strongly consistent reads for the freshness-sensitive path where required</td>
</tr>
<tr>
<td>Customer accepts at the expiry boundary</td>
<td>Aurora makes the authoritative decision</td>
</tr>
<tr>
<td>Customer submits twice</td>
<td>Return the original result using the idempotency key</td>
</tr>
<tr>
<td>Two devices accept the same quote</td>
<td>Only one conditional update succeeds</td>
</tr>
<tr>
<td>A stale price is found</td>
<td>Refuse to create the quote and request repricing</td>
</tr>
<tr>
<td>One portfolio instrument has no price</td>
<td>Reject the entire portfolio quote under the all-or-nothing policy</td>
</tr>
<tr>
<td>One portfolio instrument is stale</td>
<td>Reject the entire portfolio quote under the all-or-nothing policy</td>
</tr>
<tr>
<td>Kafka delivers a duplicate</td>
<td>Deduplicate using instrument, source and sequence</td>
</tr>
<tr>
<td>Kafka delivers an older event</td>
<td>Reject it through a conditional write</td>
</tr>
<tr>
<td>A hot instrument overloads its partition</td>
<td>Monitor lag and throughput; provision or isolate capacity according to the partition-skew policy</td>
</tr>
<tr>
<td>Primary vendor fails</td>
<td>Validate the fallback before activating it</td>
</tr>
<tr>
<td>Aurora commits but the response is lost</td>
<td>A retry returns the already-created order</td>
</tr>
<tr>
<td>Event publication fails after commit</td>
<td>The transactional outbox publishes it later</td>
</tr>
<tr>
<td>Execution venue rejects the order</td>
<td>Record the rejection without changing quote history</td>
</tr>
</tbody></table>
<h2>Advantages of the Two-Database Design</h2>
<p>The separation provides several benefits:</p>
<ul>
<li><p>market-data volume does not directly overload the order database;</p>
</li>
<li><p>each database is optimized for its workload;</p>
</li>
<li><p>accepted prices remain independent of later market movements;</p>
</li>
<li><p>the latest-price store can be reconstructed through event replay;</p>
</li>
<li><p>Aurora provides transactional order processing;</p>
</li>
<li><p>price ingestion and order processing can scale independently;</p>
</li>
<li><p>historical business decisions remain auditable.</p>
</li>
</ul>
<p>The design also introduces costs:</p>
<ul>
<li><p>there is no transaction across the two databases;</p>
</li>
<li><p>quote reads are not an atomic multi-instrument snapshot;</p>
</li>
<li><p>operational support becomes more complex;</p>
</li>
<li><p>separate backup and recovery plans are required;</p>
</li>
<li><p>cross-domain reporting needs an analytical pipeline;</p>
</li>
<li><p>teams must understand which system owns each fact;</p>
</li>
<li><p>all-or-nothing freshness can make large portfolios unquotable when one instrument is stale;</p>
</li>
<li><p>hot instruments may create uneven MSK partition load.</p>
</li>
</ul>
<p>The solution is not to hide these limitations behind a distributed transaction. It is to define clear ownership and consistency boundaries.</p>
<h2>Final Design Principle</h2>
<p>The architecture can be summarized in one sentence:</p>
<blockquote>
<p>Amazon MSK and DynamoDB manage continuously changing market state, while Aurora PostgreSQL records the immutable quotes and transactional orders that represent customer and business commitments.</p>
</blockquote>
<p>The quote is not safe because its DynamoDB read and Aurora write are atomic, nor because the market price cannot change during quote creation. It is safe because the Quote Service validates the captured observations, records the exact values and <code>priceVersion</code> identifiers it used, and then treats the Aurora record as the immutable business commitment. DynamoDB ordering protects the latest-price state, while strongly consistent reads support the freshness decision where required.</p>
<p>This separation allows the price pipeline to remain fast and scalable without weakening the guarantees required for financial transactions.</p>
<p>The most important architectural decision is therefore not the choice between DynamoDB and PostgreSQL. It is deciding where market information ends and a business commitment begins.</p>
]]></content:encoded></item></channel></rss>