Learn Labs
9. Building Data Pipelines (Kafka Connect)

9.10 What actually breaks in production — Ch. 9 consolidated

Production failure catalog
0 rows
#SymptomRoot causeFix
1A tangle of bespoke pipelines nobody can maintain; adopting a new system takes a quarterAd hoc pipelines — one custom tool per pair of endpointsSingle integration substrate (Connect); think about the whole graph, not the immediate hop
2A DBA adds a column and every downstream app breaks (or all must deploy together)Loss of metadata — no schema propagation or evolutionSchema Registry + a converter that carries schemas
3Downstream team needs a field the pipeline dropped years ago; historical data unrecoverableETL over-processing; "historical data will require reprocessing (assuming it is available)"ELT-lean: only transform what benefits every consumer
4Pipeline requires constant changes as downstream needs shiftExtreme processing coupling downstream systems to pipeline-time decisionsPreserve raw data; let Streams apps decide
5Credentials leaked from a Connect config fileConnector configs contain credentials in plaintextExternal secret config providers (Vault / AWS / Azure)
6Bad records discovered days later; can't reprocessRetention shorter than bug-discovery latency"Kafka can be configured to store all events for long periods" — size retention to your detection latency
7Connect worker OOM / broker page-cache contentionConnect running on the broker machines"run Connect on SEPARATE SERVERS from your Kafka brokers"
8Connector plug-in not foundDependencies placed at the top level of plugin.path instead of in a per-connector subdirectoryOne subdirectory per connector, containing the jar and all its dependencies
9Bizarre NoSuchMethodError / version conflictsConnectors added to the Kafka Connect classpath, bringing a dependency that conflicts with Kafka'sUse plugin.path, "the recommended approach"
10JDBC connector fails: "Access denied" / driver not foundDriver missing (doesn't ship with the connector for license reasons), or table permissionsDownload the MySQL driver into /opt/connectors/jdbc; check the Connect worker log
11Deletes never appear downstream; some updates missingJDBC polling CDC — scans by timestamp / incrementing PK; "relatively inefficient and at times inaccurate"Debezium (reads the binlog/WAL directly)
12Database load spikes from the pipelineJDBC connector repeatedly scanning tablesLog-based CDC
13Duplicate documents in Elasticsearch after reprocessingKafka records had null keys (JDBC doesn't populate them) and the sink generated new IDskey.ignore=true → deterministic topic+partition+offset document ID
14File-based pipeline loses dataFileStream connectors — "many limitations and NO reliability guarantees"FilePulse / FileSystem Connector / SpoolDir
15Corrupt messages halt a sink connectorNo error tolerance configurederror.tolerance → silently drop, or route to a dead letter queue
16Syslog connector randomly stops receiving dataConnector needs to listen on a specific machine's port, but distributed mode may schedule tasks on any nodeStandalone mode for machine-pinned connectors
17Only one task runs despite tasks.max=10JDBC connector uses MIN(tasks.max, number_of_tables)Understand the connector's own splitting logic
18Connector work distributed unevenlyConnectors and tasks "may start on any node"Worker rebalancing handles it; inspect via REST API
19Source connector reprocesses everything after a crashLogical offsets not stored / mis-designedThe framework stores offsets after broker ack — a connector authoring bug if it recurs
20A single-partition offset/config topic became a bottleneck or lost orderingInternal topics misconfigured1 partition + 3 replicas + compaction (Ch. 5 §4.2)
21Consumers can't parse Connect outputJSON converter's schemas.enable mismatch, or Avro registry URL missingPrefix converter params correctly (key.converter. / value.converter.)
22Headers not visible in console consumerRequires Apache Kafka 2.7+ and --property print.headers=trueUpgrade / add the flag
23Sink writes overwhelm the target systemNo back pressureSink context provides back-pressure methods — a well-written connector uses them
24Pipeline can't guarantee exactly-once into a databaseKafka transactions can't span systems (Ch. 8)Use the sink context's external offset storage — commit data + offsets in the target's transaction