Recently I laid out, in full, every module a real internet system touches from coding to production, as a blueprint for my own PC + WSL2 homelab. On tech choices, I deliberately mirrored the stack I already know well from the Graphene platform at work, and also turned it into a concrete toy project plan that can actually run. This post pulls both pieces together β the panoramic map and the implementation plan β as a record of this stretch of my learning path.
The overall request path: building a global mental model first
The most basic path goes: the user's request hits a CDN, then a load balancer (LB), then an API gateway, which routes into the microservice cluster; the microservices talk to caches, databases, and message queues, while logs, metrics, and traces get collected centrally and sent to a monitoring/alerting system. The Graphene platform I work with at my job is essentially the enterprise-grade implementation of exactly this pattern, just at a much larger scale and with heavier governance β once this simplest possible diagram is burned into memory, everything below is just adding detail layer by layer.
1. Development & delivery: from code to production
Version control is Git, already in use locally. Code hosting and CI triggers can be GitHub / GitLab, or a lighter self-hosted Gitea. CI/CD pipeline: Jenkins, or GitLab CI. Build tool: Maven, standard on the Java side. Image registry: Harbor (I've already built sync scripts for it before, so it fits neatly), or Docker Hub for simpler cases. Artifact repository (for Jars/dependencies): Nexus or Artifactory as a later add-on. Code quality scanning: SonarQube for static analysis.
2. Ingress layer: the first stop for a user request
DNS and CDN both use Cloudflare β I already have hands-on experience with both from the Gyuba-chan site project, so the concepts carry straight over. Load balancing (LB) locally is most directly practiced with a self-hosted Nginx; in the cloud this maps to ALB/NLB. The API gateway handles routing, auth, and rate limiting at the entry point, using Spring Cloud Gateway β a component that's already part of the Graphene stack.
3. Application / microservice layer
The business services themselves run on Spring Boot. Service-to-service calls go through Feign (Spring Cloud ecosystem). Service registration and discovery uses K8s Service + CoreDNS's fixed domain names directly, without bringing in Nacos for discovery β this is the standard approach in a native K8s environment, and it's exactly how Graphene handles it too. Nacos only comes in for configuration management, and in a K8s environment it plays that one role alone: centrally managing config across environments. Circuit breaking and rate limiting use Sentinel, part of the Spring Cloud Alibaba ecosystem, which pairs naturally with Nacos. Data access uses MyBatis as the ORM. The frontend uses React β an explicit skill requirement in job postings, and also exactly the frontend experience already accumulated on the Gyuba-chan site.
4. Data storage layer
Core business data lives in MySQL. Sharding for large data volumes uses ShardingSphere, mapping to the Sharding Ops concept already familiar from Graphene. Caching is two-tiered: Redis as distributed cache, Caffeine as local cache, forming a multi-level caching architecture to speed up reads and reduce database load. Files, images, and other unstructured data go into object storage, with AWS S3 as the choice β which also happens to be exactly what's being studied right now for SAA-C03.
5. Async & messaging layer
The message queue is Kafka, for shaving peaks and decoupling asynchronously. Dead-letter/delay queues don't need a separate new component β Kafka's own mechanisms plus business-layer implementation handle failed messages and scheduled tasks. Batch scheduling maps to the renewal batch processing already familiar from work; XXL-Job or Spring Batch are options for a later stage, and it's fine to skip installing them in the first pass.
6. Search & logging layer (the EFK stack)
Full-text search and complex queries use Elasticsearch. Log collection from each service uses Filebeat. The log processing pipeline, Logstash, is optional β if Filebeat ships logs straight into ES, this layer can be skipped for now. Log querying and visualization uses Kibana. Elasticsearch + Filebeat + Kibana together are usually called the EFK stack; add Logstash and it becomes the more classic ELK stack. What's actually planned here is EFK.
7. Container orchestration & cloud infrastructure
Application packaging uses Docker. Scheduling, scaling, and self-healing are handled by K8s, already set up locally with k3d. The K8s cluster's traffic entry point is nginx-ingress, which is the next piece planned for hands-on setup. Infrastructure-as-code (IaC) uses Terraform or Ansible, part of the long-term plan. The cloud platform underpinning all of this is AWS, matching the SAA-C03 material currently being studied.
8. Observability's three pillars
Observability is usually split into three pillars. Metrics are collected and stored by Prometheus, with Grafana turning them into dashboards β both already underway. Alerting is handled by Alertmanager, built into the Prometheus ecosystem, firing notifications on threshold breaches. The logging pillar is the EFK stack described above. Tracing uses Jaeger or Zipkin, a more advanced piece that lets you see how a single request propagates across the gateway, service A, and service B β and incidentally, the "Sleuth + Zipkin" combo on the Graphene skills sheet is exactly the classic tracing pairing in the Spring ecosystem; once basic monitoring is running smoothly, this layer is well worth adding.
9. Security & networking (advanced, can wait)
User authentication and authorization use Spring Security + JWT. Secrets management starts with K8s Secret, with Vault as a later option. Service-to-service access control uses K8s NetworkPolicy. This whole section is lower priority overall β to be circled back to once the earlier skeleton is running smoothly.
The seven build stages the panoramic map lays out
Mapping the whole panoramic diagram onto my own PC produces seven stages. Stage 1 is the K8s base (k3d) plus Prometheus/Grafana monitoring, already begun. Stage 2 is writing two simple Spring Boot microservices, with inter-service calls going directly through K8s Service DNS rather than Nacos β Nacos is limited to the single role of "config center" here. Getting the path "gateway routing β K8s DNS call β config pulled from Nacos" working end-to-end is the foundation for all microservice practice, and matches Graphene's real architecture. Stage 3 wires in MySQL, Redis/Caffeine, and MyBatis, to experience real phenomena like cache penetration and cache breakdown firsthand. Stage 4 adds Kafka, building an "order β async stock deduction" style scenario to experience message buildup and consumer lag. Stage 5 builds the EFK log pipeline, centralizing logs from the earlier services. Stage 6 sets up Jenkins CI/CD, closing the loop of "code push β auto build β auto deploy to k3d." Stage 7 is advanced work: custom Ingress, tracing (Zipkin), and managing the whole infrastructure with Terraform.
Advanced challenge: recreating Graphene's Zorro BLCS β the in-house CDC middleware
Zorro BLCS's core mechanism is disguising itself as a MySQL replica, tailing the primary's binlog, wrapping change events uniformly into a ZorroEvent format, and pushing them to a Kafka cluster β downstream consumers then split by topic, each consuming independently with no awareness of one another. Two consumer paths handle this completely differently: the CDC microservice parses ZorroEvents and inserts/updates/deletes documents in Elasticsearch, so queries like the policy list can go through the ES-based CQRS read path; the Alphad microservice instead reconstructs the ZorroEvent back into SQL and executes it directly, mirroring the source database's writes verbatim into the CBQ reporting database β this "binlog event β reconstructed SQL β replayed on the target DB" pattern is essentially an in-house logical replication, functionally similar to MySQL's own row-based binlog replication, just carried out across systems via Kafka. Key design points confirmed from BLCS-OPS screenshots: the position uses a MySQL GTID Set, formatted as server-uuid plus transaction sequence β GTID's advantage is staying globally unique and monotonically increasing across primary/replica failover, so a restarted task can resume from its last position; there's an "auto-replay missed messages" option, letting binlog events lost during an outage be selectively replayed on restart, a backstop against data loss; report and cdc are two independently configured BLCS jobs, each connecting to its own read replica and publishing to its own Kafka topic; filter dimensions include database-name expressions, table-name regex (matching the numeric suffixes from sharding), and the sharding key itself (for downstream cross-shard aggregation); the source connection uses a dedicated read-only account with an _ro suffix, isolating CDC reads from impacting the production primary's write path; and there's a "use zorro format" option deciding whether binlog changes matched by a given filter rule get wrapped in the unified ZorroEvent protocol.
The plan is to recreate this in the homelab using the open-source Canal β mechanically identical to Zorro BLCS β to simulate the production side: connecting to my own MySQL, subscribing to binlog, publishing to Kafka. On the consumer side, two simple services: one simulating the CDC microservice (consuming and writing to a local Elasticsearch), and one simulating the Alphad microservice (consuming, reconstructing events into SQL, and executing them against a different MySQL schema to simulate the reporting-DB mirror). This should give a full hands-on feel for GTID positions, resuming from checkpoints, and "the same change event, two completely different consumption strategies." Two things are already confirmed and don't need re-checking with colleagues: the report topic is consumed by the Alphad microservice specifically to do SQL mirror execution; and Zorro is a unified CDC event protocol plus tool name, not merely a format option. Still-open questions are the exact schema/field structure of ZorroEvent, and whether the "assign machine" button is itself a kind of service-discovery/scheduling mechanism comparable to the ZK approach used for batch scheduling.
ES + CQRS: read/write separation, with reads leaning hard on ES
After a microservice writes to MySQL, it actively publishes a Kafka message; an independent CDC consumer service subscribes and writes into an ES index; queries like the policy list read ES directly and never touch the database β this is the standard CQRS (Command Query Responsibility Segregation) pattern: the write path goes through MySQL, the read path through ES, with async messaging bridging the two toward eventual consistency. Common failure points: the CDC consumer service lagging or crashing, causing ES data to fall behind the real database so data a user just created "temporarily can't be seen"; and Kafka messages being lost or double-consumed, causing missing or duplicate data in the ES index. Also worth checking: is there a scheduled full comparison between the DB and ES as a compensating mechanism, and is there retry-on-failure or a dead-letter queue? The homelab plan is a simple Spring Boot service that writes MySQL then publishes a Kafka message, with another service consuming it and writing to Elasticsearch β a small demo where list queries go through ES and detail queries through MySQL β then deliberately killing the consumer for a while to observe what the data-inconsistency window actually looks like in practice.
In-house batch scheduling + Zookeeper: strong-consistency coordination vs. K8s-native discovery
The batch job code itself is scattered across the various microservice codebases, while an independent scheduling service does service discovery through Zookeeper β each microservice registers to an ephemeral znode in ZK rather than going through K8s Service. The technical judgment behind this: a K8s Service can only answer "route to one healthy instance," not questions like "what instances currently exist, and who's the leader," which need an exact instance list. Common batch-scheduling needs β sharded scheduling (hashing task IDs across multiple workers) and leader election (only one of several scheduler instances should actually dispatch tasks) β inherently require CP-style strong consistency and distributed coordination primitives like ZK's ephemeral nodes, sequential nodes, and watch mechanisms, not just service-level load balancing. The more likely real story: the batch scheduling framework predates the system's full move to K8s, and ZK is simply baked into the framework's own logic (possibly built on an open-source framework like elastic-job/tbschedule, or an in-house variant with the same idea) β a natural result of legacy history plus framework lock-in, rather than the team deliberately choosing two separate discovery mechanisms. Open questions still pending confirmation: whether the batch scheduling framework is homegrown or built on top of an open-source one, and whether the ZK cluster has any history of stability issues like connection flakiness or session-timeout misconfiguration. The homelab plan for this is a single-node Zookeeper locally, writing a "multi-instance registration + leader election" demo using Curator (ZK's officially recommended Java client) to get hands-on with ephemeral nodes and watches, then writing an equivalent pure-K8s-Service-DNS version for comparison, to feel the fundamental difference between the two.
More modern additions: the real gaps in an otherwise solid skeleton
Gateway plus microservices plus K8s plus CDC/CQRS plus monitoring and logging is already a textbook-standard scaffold for a mid-to-large company, but from a 2026 vantage point there are still a few genuine gaps. First, distributed transactions: CQRS solves eventual consistency under read/write separation, but a more fundamental problem β how to guarantee that writes spanning multiple microservices in one business operation either all succeed or all roll back β hasn't been touched yet. In a monolith, one @Transactional handles it; once split into microservices there's no built-in answer, and the standard solutions are the Saga pattern (each failed step triggers compensating actions for the earlier ones) and TCC (Try-Confirm-Cancel, in three phases). The homelab plan is to wire the open-source Seata (from Alibaba) into the toy order service, deliberately fail a step, and watch how the compensation mechanism undoes the earlier operations. Second, canary releases / blue-green deployment: the production-standard approach is validating a new version against a slice of traffic before switching over fully; K8s natively can approximate this with weighted Service routing, and Argo Rollouts (more advanced) can capture the gap between "can be released" and "dares to be released." Third, load testing and chaos engineering, especially important for the SRE direction: load testing with k6/JMeter/Gatling to find peak QPS and performance bottlenecks; chaos engineering with Alibaba's open-source ChaosBlade or Chaos Mesh, actively injecting faults (killing pods, simulating network latency, simulating connection-pool exhaustion) to see whether the system self-heals and whether alerts fire in time β the plan is to load-test with k6 first, then randomly kill pods with Chaos Mesh, turning RCA skills from passively waiting for incidents into something actively practiced. Fourth, Service Mesh: the Spring Cloud Gateway + Feign combo requires manually wiring circuit-breaking, retries, timeouts, and encryption into business code, whereas a Service Mesh like Istio pushes these capabilities down into a sidecar proxy at the infrastructure layer, completely invisible to business code β not that Spring Cloud is obsolete, but it's worth experiencing a more modern style of service governance. Fifth, an OLAP reporting engine: mirroring Alphad's output into MySQL works, but MySQL is fundamentally a row store optimized for transactions, and massive aggregation queries run much slower than on a columnar store; the modern approach leans toward dedicated analytics engines like ClickHouse or Apache Doris, directly connected to the big-data platform discussed next, so it's planned to be practiced together with it.
The big-data platform: from single transactions to mass-scale analysis
The grandfather of big-data platforms is Hadoop, whose standard trio from roughly 2006β2013 was HDFS (a distributed filesystem managing where data lives), MapReduce (a compute framework managing how data gets processed, though every step spills to disk, making it fairly slow), and YARN (resource scheduling, managing which machine in the cluster a task lands on). The modern big-data stack is essentially the result of these three components being replaced one by one: HDFS gives way to MinIO/S3, lighter and more cloud-native; MapReduce gives way to Spark, memory-based and an order of magnitude faster; YARN gives way to K8s, absorbed uniformly into container orchestration. The skeletal logic was invented by Hadoop, but the concrete implementation has since moved on β the homelab plan doesn't spend time deploying Hadoop itself, but it's worth recognizing this terminology in case an older system still running Hadoop/Hive turns up at work someday. Mapped onto concrete layers: the data-ingestion layer reuses the existing Canal (simulating Zorro BLCS) plus Kafka, adding one more consumer to the same binlog events; the data-lake storage layer uses MinIO, open-source and S3-protocol-compatible, the best fit for a home lab; the data-warehouse/OLAP layer uses ClickHouse or Apache Doris, columnar and optimized for aggregation queries; the batch compute engine is Spark, the industry standard, especially Spark SQL; the real-time stream compute engine is Flink, which can directly consume CDC events from Kafka for real-time aggregation, such as real-time premium totals; task scheduling uses Airflow, the industry standard with DAG visualization; metadata management/data catalog uses DataHub, an optional advanced piece that can wait; and BI/visualization uses Superset, an open-source Apache Foundation project that's quick to pick up.
The most valuable part here is the connection point with the existing architecture: there's no need to manufacture data from scratch β the Kafka + Canal CDC pipeline built in Stage 4 can be reused directly, with the same binlog change events feeding, alongside the existing ES-writing consumer, one more consumer that writes into the data lake / triggers Flink real-time computation. This gives a genuine feel for the perspective shift where "the same data looks like part of CQRS from the business system's view, but an analytics data source from the big-data platform's view." The suggested build order: first MinIO, to get data actually landing somewhere with the least effort; then ClickHouse, switching the reporting data previously mirrored by Alphad from MySQL over to ClickHouse, for a direct, intuitive comparison of query-speed differences; then Flink, consuming CDC events from Kafka to build a small rolling real-time premium-total demo, writing into ClickHouse; then Superset, connecting to ClickHouse and dragging out a dashboard in five minutes β probably the most satisfying step of all; and finally, as an advanced add-on, Airflow, orchestrating a full nightly reconciliation/aggregation job, which can be compared against Graphene's in-house batch scheduling to feel the design differences between data-ETL scheduling and business-batch scheduling. Spark comes last, since the batch engine overlaps conceptually with stream processing, and will make more sense once stream processing is already familiar.
AI and vector databases: a new module explained from scratch
This area hasn't been touched before at all, so it's worth starting from the most basic concepts. What's a vector/embedding: feed a piece of text or an image to an embedding model, and it outputs a string of numbers (say, 1536 floating-point values) representing that text's semantic meaning β the key property is that texts with similar meaning produce number strings that are also mathematically close, measurable with cosine similarity. For instance, "insurance claim" and "claim application" look completely different on the surface but end up mathematically close once converted to vectors. What's a vector database: a traditional database like MySQL is good at exact matching, while a vector database is good at similarity search β a completely different indexing structure and query style. RAG (retrieval-augmented generation) is currently the most central pattern in AI applications, solving the problem that an LLM doesn't know private data and its knowledge goes stale: private documents (say, the SLI-BAT001 spec docs already studied in depth) get chunked into small pieces, each converted to a vector and stored in a vector DB; when a user asks a question, that question is also converted to a vector, used to retrieve the most similar chunks from the vector DB; those chunks plus the user's question are then handed to the LLM, asking it to answer using that material as reference β so the LLM's answer draws on real private material rather than hallucinating from training memory, which is also the most mainstream fix for the LLM-knowledge-goes-stale problem.
New modules the architecture needs include: an embedding model, running locally via Ollama (e.g. nomic-embed-text) rather than calling a paid API; a vector database, starting with pgvector for the homelab β a PostgreSQL extension, which is ideal since a relational database is already familiar territory and installing a plugin gives vector capability without learning a whole new system, with Milvus as a more full-featured but more complex option for later; an LLM inference service, running an open-source model locally via Ollama; an orchestration framework, which decides when to call the LLM, whether to retrieve material first, whether to call tools, and how to handle results β standardizing this "glue logic" matters a lot, and LangChain suits chain-like/linear flows (RAG's "fetch material β feed the LLM β produce a result" fits well), while LangGraph suits graph-shaped flows with loops or conditional branching, and is currently the mainstream framework for building agents; and an AI gateway, optional and more advanced, using LiteLLM Proxy to unify rate limiting, routing, and cost tracking across multiple LLM calls. RAG can be hand-coded without an orchestration framework, but once the flow gets more complex β first classifying which knowledge base a question belongs to, then deciding whether to retrieve, possibly calling external tools like a calculator or weather lookup afterward, and finally summarizing β hand-rolled if-else quickly spirals out of control. LangChain packages common steps (loading documents, chunking, embedding, retrieval, prompt assembly, calling the LLM, parsing output) into composable "chains," while LangGraph models the whole flow as a state machine supporting loops, conditional jumps, and "retry until satisfied." Agents β an LLM autonomously deciding what to call next and whether to retry β are now mostly built on LangGraph or similar frameworks across the industry; the two aren't either/or, and many real projects use LangChain for the simple RAG portion and LangGraph for more complex multi-step agent logic.
The suggested build order: first PostgreSQL + the pgvector extension, far simpler than standing up a dedicated vector database; then run a small embedding model via Ollama and write a script to chunk BSD study documents (e.g. SLI-BAT001) by paragraph, convert them to vectors, and store them in pgvector; then write a simple retrieval demo β typing in a question like "what's the process for endorsement changes" and retrieving the most relevant chunks from pgvector, just to verify that the semantic-similarity idea really works; then connect Ollama's chat model, feeding it the retrieved passages plus the question, running a full RAG flow end to end β once this step is done, there's a private AI assistant that can answer BSD-spec questions, directly useful for current BSD studies; later, once the PC has a 3090, this same pipeline can drop in a larger local model. This area also feeds back into two existing projects: for BSD study, all the deeply-studied SLI-API/BAT/ACC/USR documents can be fed in to build a dedicated insurance-domain Q&A assistant; for the Gyuba-chan family, if the site ever wants semantic search over story content or simple character Q&A interactions, this RAG experience carries over directly.
From blueprint to implementation: the Mini Policy System
The panoramic map alone isn't enough β a concrete, runnable implementation plan is needed, one Claude Code can build to spec inside VSCode. There are five design principles. First, unify the business domain and avoid using middleware just for the sake of using it β all services revolve around the same simplified domain, a mini policy management system, with naming and fields drawn from the already-familiar TENGAN/INGURAMU-style terminology to make the practice feel more grounded. Second, one codebase, two environments: local development spins up MySQL/Redis/Kafka/ES etc. with a single docker-compose command, while the services themselves run directly via IDE or Claude Code (not containerized, for easier debugging); the "production" environment runs everything β services and middleware alike β on k3s, deployed through Harbor and a CI/CD pipeline, simulating a real release process. Third, producer/consumer relationships must be genuine, not decorative β at least two independent Kafka consumer paths are required, mirroring the panoramic map's "one event, two consumption strategies" design. Fourth, every stage must actually run and show visible results β no piling up middleware that's installed but never used. Fifth, leave hooks for the later big-data/AI stages without implementing them in this first phase β just avoiding self-inflicted design traps, e.g. making sure the Kafka topic design and event schema anticipate being reused by multiple future consumers.
The business domain in one sentence: a user logs in, creates/queries policies, policy creation asynchronously triggers a notification and an async write to the search index, and policies can then be searched. The core Policy entity has id, policyNo (policy number), holderName, productType (TENGAN or INGURAMU, mapping to the two familiar product lines), premium, status (DRAFT/ACTIVE/CANCELLED), and created/updated timestamps. There are five services in total: frontend, a React + Vite SPA handling login, the policy list/creation form, and the search page; gateway-service, the API gateway on Spring Cloud Gateway, handling routing, JWT auth, and rate limiting; policy-service, the core write service on Spring Boot + MyBatis + Liquibase + MySQL + Redis, handling policy CRUD on the write path, Redis-cached detail lookups, publishing Kafka events on create/update, and auto-migrating tables via Liquibase on startup; notification-service, consumer A, a pure consumer with no database, subscribing to the policy-events topic to simulate sending notifications β just logging or writing to a local file is enough; and search-service, consumer B plus a read service on Spring Boot + Elasticsearch, subscribing to the same topic to write into ES and exposing a read-only policy search API, forming the CQRS read path. Table-structure management skips the crude approach of detecting-and-creating tables in application code at startup, instead using Liquibase, matching the tool actually used at the company β changesets numbered sequentially under a changelog directory, with the application checking the DB's current version on startup and applying whatever changesets haven't run yet, making schema creation itself version-controlled and traceable.
The directory structure is a single monorepo, everything in one folder for unified management in VSCode: frontend, gateway-service, policy-service, notification-service, and search-service each get their own directory. Under infra sits docker-compose.dev.yml (local dev middleware), a k8s subdirectory (one deployment-manifest subdirectory per service, plus a middleware subdirectory for the production-facing middleware running in k3s), and a jenkins subdirectory (Jenkinsfile templates). A docs directory holds kafka-event-schema.md as the shared event-format contract across services. The root has a single README as the project overview. Each service's subdirectory is required to have its own README, clearly stating what the service does, how to run it standalone locally, which middleware it depends on, and which ports/APIs it exposes.
The environment setup uses a three-way split rather than a simple local/production dichotomy. The local dev docker-compose installs only middleware (MySQL, Redis, Kafka in KRaft mode so no separate Zookeeper is needed, Elasticsearch, Kibana), while the services themselves run directly via mvn spring-boot:run or npm run dev for easy breakpoint debugging. The core decision on the "production" side is that stateful middleware β MySQL and Redis β doesn't go into k3s at all; instead it runs on the host machine as long-running Docker containers managed by systemd, playing the role of a cloud vendor's managed service (RDS/ElastiCache). The reasoning: production environments don't typically manage a stateful DB's high availability inside the business cluster itself, so simulating this external dependency relationship is closer to real architecture than building a simplified HA setup inside k3s β and having MySQL run on the host actually makes the later Canal binlog-watching more direct. For the k3s business cluster to reach MySQL/Redis on the host, a new concept comes in: the "Service without a selector." A normal Service uses a selector to automatically associate with matching-labeled Pods in the same namespace; a selectorless Service is specifically for pointing at resources outside the cluster β a manually written Endpoints object points at the host's IP and port, while business code still sees a standard Service DNS name and never has to know the database is actually outside the cluster β the same decoupling idea as a real cloud app not caring where RDS is physically deployed. Kafka and Elasticsearch, being either stateless or self-redundant by design, are fine to run inside k3s, deployed as single-replica instances via the Bitnami Helm chart, with the focus on practicing writing YAML/Helm values rather than obsessing over production-grade multi-replica configuration. On image flow: business-service images all get pushed to Harbor, and k3s pulls from Harbor; images built locally need a docker push to Harbor, with k3s configured with imagePullSecrets pointing at it β a good opportunity to practice private-registry authentication hands-on. The whole system gets its own toy-system namespace, separate from k3s's built-in kube-system, for easier management and later full cleanup. There's also one independent, non-blocking practice module: deploying an extra single-node MySQL StatefulSet (called mysql-statefulset-demo), with its own test database not connected to any business service, purely to experience that data survives Pod deletion/recreation, the PVC-PV binding relationship, and how this fundamentally differs from the selectorless-Service approach β slotted in any time after Stage 5 and before Stage 6.
CI/CD pipeline component choices: the code repository starts with a self-hosted Gitea β much lighter than GitLab CE and better suited to a machine with, say, 8 cores and 15GB of RAM; it's swappable for GitLab CE later if practicing more enterprise-like features (MR approval flows, a CI variable-management UI) becomes worthwhile. The CI/CD engine is Jenkins, already planned. The image registry reuses the existing Harbor experience. The first-phase deployment method is a Jenkins pipeline running kubectl set image / kubectl apply directly β simple and direct, prioritizing getting the loop closed. A second, more advanced phase adds Argo CD for GitOps, aligned with the canary-release direction and a more modern deployment paradigm. Taking policy-service as an example, the pipeline stages are: a developer's git push hits Gitea, Gitea's webhook triggers Jenkins, and the Jenkins pipeline runs, in order, checkout code, mvn test for unit tests, mvn package to build the Jar, docker build with a tag, docker push to Harbor, and finally kubectl set image to trigger a K8s rolling deployment β a good chance to watch the Deployment's rolling-update process along the way. The Jenkinsfile template gets written as a parameterized template, with each service's own Jenkinsfile just referencing it and passing service-specific parameters like name and port, cutting down on duplication.
The phased implementation order aligns with the panoramic map but is re-sequenced for this toy system, across twelve stages in total. P0 is already done: a single-node k3s plus Docker. P1: bring up MySQL/Redis locally via docker-compose and write policy-service (including Redis caching), getting CRUD working locally. P2: add gateway-service with JWT auth, and add the frontend's login and list page, wiring up the full path from frontend through gateway auth to backend. P3: add Kafka locally, have policy-service publish events on policy creation, and write notification-service to consume them, experiencing the decoupling of "one event, independent consumers." P4: add ES locally, write search-service to consume the same event and write to ES while exposing a search API, add a search page on the frontend, completing the full CQRS read/write-separation path. P5: bring up MySQL/Redis on the host as long-running Docker containers simulating cloud-managed services, create a toy-system namespace in k3s, connect business services to the host's MySQL/Redis via selectorless Services, deploy Kafka/ES as single-replica Helm releases into k3s, and manually kubectl apply all business services onto k3s (no CI/CD yet) β running the full path once in the "production" environment. P5.5: an inserted practice module, an independent MySQL StatefulSet + PV demo not connected to business data. P6: only now install Prometheus/Grafana, since there's real traffic and services to observe β real CPU/memory/QPS curves visible, and a Pod can be deliberately killed to watch self-healing. P7: wire in Gitea + Jenkins + Harbor, turning P5's manual deployment into an automated pipeline, so a git push triggers a deploy and rolling releases work end to end. P8: wire the EFK log pipeline into these services, so troubleshooting means searching logs rather than checking kubectl logs one by one. P9 (advanced): introduce Canal watching the host MySQL's binlog, adding an independent report-events topic and report-service β Canal publishes to report-events, and report-service consumes it, reconstructs SQL, and mirrors it into ClickHouse rather than MySQL, experiencing OLAP columnar storage's aggregation advantage; the existing policy-service manually publishing policy-events stays unchanged, letting the two pipelines run side by side for comparison, fully recreating Zorro BLCS's design idea of "one binlog, two independent consumption pipelines." P10 (advanced): Argo CD for GitOps-style deployment plus a simple canary-release demo. P11 (advanced): load-testing policy-service with k6 and randomly killing Pods with Chaos Mesh. Big data and AI/RAG are treated as independent follow-on extensions, not crammed into this first phase, but already compatible by design β the policy-events topic can gain another future consumer writing to MinIO or triggering Flink without affecting the two existing consumers, and the policy data in ES can also later serve as a data source for pgvector/embedding experiments.
The Kafka event contract is written in docs/kafka-event-schema.md, a part Claude Code needs to follow strictly. The event structure for the policy-events topic includes eventId, eventType (POLICY_CREATED/POLICY_UPDATED/POLICY_CANCELLED), an occurredAt timestamp, and a nested policy object (id, policyNo, holderName, productType, premium, status) β a single topic carrying multiple event types, split by the eventType field, matching most Kafka event-design practice. The report-events topic added in Stage P9, produced by Canal watching the host MySQL's binlog, aligns its format with the ZorroEvent idea β wrapping binlog changes in a unified event protocol rather than reusing the policy-events application-layer schema β and includes gtid, database, table, eventType (INSERT/UPDATE/DELETE), data (the changed row data), and occurredAt. report-service consumes this topic, reconstructs the events into SQL, and replays them in ClickHouse, mirroring Alphad's replication pattern; this pipeline is a completely independent topic from policy-events, with no effect on each other.
There are eight implementation requirements for Claude Code. Each service gets its own independent pom.xml/package.json rather than a parent-child module structure, keeping things simple. All service configuration is injected via environment variables, never hardcoded β locally via .env or an IDE run configuration, and via ConfigMap/Secret inside k3s. Each service exposes /actuator/health, preparing for the later Prometheus integration. The code in policy-service that publishes Kafka events gets encapsulated into its own EventPublisher class, so that when it's swapped for Canal in Stage P9, only the layer deciding "who triggers the event send" needs to change, while EventPublisher itself and the downstream consumers stay untouched. Implementation proceeds in P1-through-P4 order, with each stage expected to run and show visible results before moving on, rather than writing all the service code up front and only then integrating. Table-structure changes in policy-service go exclusively through Liquibase changelogs β hand-written "check if table exists, create if not" logic in Java is not allowed. The selectorless-Service-plus-Endpoints configuration for Stage P5 gets written out clearly in its own file, with a comment at the top explaining that this simulates a cloud-managed database pointing at the host's IP, so the intent is obvious at a glance when revisiting it later.
Closing thoughts: separating "understanding the architecture" from "writing business code"
There's a deliberate division of labor behind this whole plan: what's being practiced is the architectural skeleton and how the components coordinate with each other; the business-logic code itself β the toy order service, the embedding-chunking script, the Flink aggregation job, and so on β doesn't need to be hand-written from scratch, and handing it to Claude Code to implement saves time, leaving focus on the core goal of understanding what each layer does, why it's designed that way, and how to troubleshoot it. The panoramic map is about knowing what exists; the implementation plan is about actually running it by hand. Together, the two roughly amount to the homework assigned to myself for this stretch of time.