QuestionHow do you efficiently store and query time-stamped data?

Time-Series Database Architecture: Storing Millions of Metrics

  • medium level
  • 58-sec video
  • 5-min read
  • 5-question quiz

The big idea

Time-series databases are designed to store and query time-stamped metrics efficiently.

Your path

Log in to save progress

Explain it like I’m 10

Imagine a big notebook where you write down the temperature every hour. A time-series database is like that notebook but can handle thousands of temperatures from many places at once!

Introduction to Time-Series Databases

Time-Series Databases (TSDB) are specialized databases designed to handle time-stamped data collected over time. They are crucial for applications that monitor metrics like CPU utilization, temperature, or network traffic from various sources.

Key Components of Time-Series Database Architecture

  1. Time-Series Data: This consists of timestamps, values, and optional labels. For example, a CPU utilization metric might look like this: { "timestamp": "2023-10-01T12:00:00Z", "value": 75, "label": "server1" }.
  2. Ingestion Layer: This is where metrics flow from servers into the database. Efficient ingestion is critical to handle high volumes of data.
  3. Storage Architecture: Measurements are organized for time-based access, often in a way that allows for quick retrieval of recent data.
  4. Time-Range Queries: Users can query metrics over specific time windows, such as the last 15 minutes, to analyze trends or anomalies.
  5. Compression: To save storage space, time-series databases often use compression techniques to reduce the overhead of storing large volumes of measurements.
  6. Retention Policies: These define how long historical data is kept. For example, you might keep high-resolution data for a week and lower-resolution data for a year.
  7. High Cardinality: This refers to having too many distinct label combinations, which can increase resource usage and complicate queries.

Example Use Case

Consider a scenario where you are monitoring CPU utilization across thousands of servers. The data flow would look like this:

Servers → Ingestion Layer → Time-Series Database → Query & Aggregation → Dashboard

Trade-offs in System Design

When designing a time-series database, you must balance:

  • Ingestion rate
  • Storage capacity
  • Retention of historical data
  • Query performance

Common Interview Trap

High data volume and high cardinality are not the same. Uncontrolled label cardinality can lead to a vast number of distinct time series, increasing both storage and query overhead.

Memory Trick

INGEST FAST — Receive telemetry efficiently
STORE BY TIME — Organize measurements for time-oriented access
QUERY THE WINDOW — Retrieve and aggregate the time range you need

Conclusion

Designing a time-series database requires careful consideration of ingestion throughput, time-range queries, compression, retention, and cardinality. The exact strategies for storage and indexing will depend on the specific implementation.

Where you’ll see this

Time-series databases are used in monitoring applications, like tracking server performance or weather data.

How exams test this

Exams may test your understanding of time-series architecture and its components, often trapping students with the difference between data volume and cardinality.

📖 Words to know

Time-Series Data
Data that is indexed by time.
Ingestion Layer
The component that collects and inputs data into the database.
Cardinality
The uniqueness of data values in a dataset.

Got it? Lock it in 🔒

5 quick questions. Students who test themselves remember far more than those who just re-read.

Learn next

Watch first

Course outline · 268 topics
  1. 1What FAANG Interviewers Evaluate in System Design Interviews
  2. 2Understanding High-Level Design (HLD) and Low-Level Design (LLD)
  3. 3Understanding Functional and Non-Functional Requirements
  4. 4Understanding Scalable System Design in FAANG Companies
  5. 5Understanding Latency, Throughput, and Response Time
  6. 6Understanding Availability vs Reliability in System Design
  7. 7Understanding Latency and Throughput in System Design
  8. 8Understanding Vertical and Horizontal Scaling in System Design
  9. 9Stateless vs Stateful Servers: Scaling Made Easy
  10. 10Monolith vs Microservices: Choosing the Right Architecture
  11. 11Understanding Load Balancers in System Design
  12. 12L4 vs L7 Load Balancers: Key Differences Explained
  13. 13Understanding Reverse Proxy in High-Level Design
  14. 14Understanding CDN: How Cache Hits Improve Performance
  15. 15Understanding Cache Hits: Reducing Database Load
  16. 16Understanding the Cache-Aside Pattern in System Design
  17. 17Redis: Why Is It So Fast for High-Scale Systems?
  18. 18SQL vs NoSQL: Choosing the Right Database for Your Needs
  19. 19Understanding Database Replication in High-Level Design
  20. 20Understanding Sharding: How Databases Manage Large Data
  21. 21FAANG HLD 🔥 | Partition Pruning Explained — Scan Less, Query Faster! ⚡
  22. 22B-Tree Index: Efficient Database Navigation
  23. 23Read-After-Write Consistency in Database Systems
  24. 24Understanding Read Replicas in Database Scaling
  25. 25Identifying and Optimizing Database Bottlenecks
  26. 26Understanding CAP Theorem: Trade-offs in Distributed Systems
  27. 27Understanding Consistency Models in Distributed Systems
  28. 28Understanding Quorum Reads in Distributed Systems
  29. 29Leader Election in Distributed Systems: Handling Leader Failures
  30. 30Log Replication and Majority Commit in Distributed Systems
  31. 31Consistent Hashing: Key to Distributed Systems Scalability
  32. 32Distributed Caching: Why Spread It Out?
  33. 33Cache Eviction Policies: LRU, LFU, and FIFO Explained
  34. 34Cache Invalidation: TTL vs Freshness Explained
  35. 35Understanding Cache Stampede in High-Level Design
  36. 36Understanding Cache Penetration in System Design
  37. 37Understanding Cache Breakdown and Request Coalescing
  38. 38Token Bucket Rate Limiting Explained for Interviews
  39. 39Token Bucket vs Leaky Bucket: Key Differences Explained
  40. 40API Gateway: The Front Door of Your System
  41. 41API Gateway vs Proxy vs Load Balancer: Key Differences
  42. 42Service Discovery in Microservices Architecture
  43. 43DNS Service Discovery: Is DNS Enough for Microservices?
  44. 44Understanding DNS Record Types: A, CNAME, MX, and More
  45. 45Understanding DNS Resolution: How URLs Load in Browsers
  46. 46Understanding TCP 3-Way Handshake: SYN, SYN-ACK, ACK
  47. 47Understanding the TLS Handshake in HTTPS Connections
  48. 48HTTP vs HTTPS: Understanding the Key Differences
  49. 49Understanding HTTP Methods and Status Codes
  50. 50REST API Design: Building Scalable APIs for Real-World Use
  51. 51API Pagination: Efficiently Handling Large Datasets
  52. 52Database Indexing: How to Make Queries Fast
  53. 53Composite Indexes: Speed Up Multi-Column Queries in Databases
  54. 54Read vs Write Scaling: How to Scale Databases
  55. 55Database Replication: Scale and Survive Failures
  56. 56Sync vs Async Replication: Choosing the Right Approach
  57. 57Failover and Split-Brain in High-Level Design
  58. 58Understanding Write-Ahead Log (WAL) in Databases
  59. 59Understanding Transactions and ACID Properties in Databases
  60. 60Understanding Database Isolation Levels in Software Engineering
  61. 61Optimistic vs Pessimistic Locking in Database Management
  62. 62Understanding 2-Phase Commit in Distributed Transactions
  63. 63Understanding the Saga Pattern in Distributed Transactions
  64. 64Transactional Outbox Pattern: Never Lose Events!
  65. 65Understanding Idempotency in System Design
  66. 66Understanding Message Queues in Distributed Systems
  67. 67Queue vs Pub/Sub: Key Differences in System Design
  68. 68Understanding Apache Kafka for Scalable Event Streaming
  69. 69Understanding Kafka Partitioning for Scalable Messaging
  70. 70Message Delivery Semantics: At-Most-Once vs At-Least-Once vs Exactly-Once
  71. 71Understanding Kafka Consumer Lag and Its Impact
  72. 72Kafka Retention vs Compaction: History or Latest State?
  73. 73Kafka Exactly-Once Processing: Avoiding Duplicates
  74. 74Kafka Schema Evolution: Changing Schemas Without Breaking Consumers
  75. 75Kafka Producer Reliability: ACKs, Retries & Idempotence
  76. 76Kafka Ordering and Keys: Ensuring Message Order
  77. 77Kafka Consumer Offsets: Auto Commit vs Manual Commit
  78. 78Understanding Kafka Retries and Dead Letter Queue (DLQ)
  79. 79FAANG HLD 🔥 | Kafka Backpressure Explained — What Happens When Consumers Can't Keep Up? 🚨
  80. 80Understanding Kafka Brokers, Controllers, and KRaft Architecture
  81. 81Kafka Replication and ISR: Ensuring Data Availability
  82. 82Kafka Partition Reassignment: Move Replicas Without Data Loss
  83. 83Kafka Consumer Group Rebalancing Explained
  84. 84Kafka Consumer Assignment Strategies: Range vs RoundRobin vs Sticky
  85. 85Kafka Consumer Liveness: Heartbeats, Sessions & Timeouts Explained
  86. 86Kafka Consumer Lag Monitoring: Prevent Production Issues
  87. 87Kafka Lag Recovery Time: How Fast Can Consumers Catch Up?
  88. 88Kafka Consumer Fetch & Batching: Boosting Throughput
  89. 89Kafka Fetch Limits vs Poll Records: Understanding the Difference
  90. 90Kafka Eager vs Cooperative Rebalancing: Key Differences
  91. 91Kafka Static Membership: Reducing Unnecessary Rebalancing
  92. 92Understanding Kafka Rebalance Listeners in High-Level Design
  93. 93Understanding Kafka Offset Reset: Earliest vs Latest
  94. 94Kafka Replay Safely: Reprocess Messages Without Losing Data
  95. 95Kafka Commit Strategies: Auto vs Manual Commit Explained
  96. 96Kafka commitSync vs commitAsync: Choosing the Right Method
  97. 97Kafka Async Commit Ordering: Can Older Offsets Overwrite Newer Ones?
  98. 98Kafka Consumer Concurrency: Understanding Partitions and Threads
  99. 99Kafka Consumer Multithreading: Processing Messages in Parallel
  100. 100Kafka Consumer Parallelism: Partitions, Keys, and Throughput
  101. 101Kafka Consumer Pause & Resume: Managing Backpressure Smartly
  102. 102Kafka Backpressure Strategies: Handling Traffic Spikes
  103. 103Kafka Graceful Shutdown: Stop Without Losing Work
  104. 104Kafka Error Handling and Retry Strategies in HLD
  105. 105Kafka Retry Topics and Delayed Retries Explained
  106. 106Understanding Kafka Dead Letter Topics and Poison Messages
  107. 107Kafka Poison Messages: Risks of Infinite Retries
  108. 108Kafka Consumer Reprocessing: Safe Event Replay Strategies
  109. 109Kafka Replay Without Breaking Production: Safe Architecture
  110. 110Kafka Pause vs Seek: Key Differences Explained
  111. 111FAANG HLD 🔥 | Kafka Seek — Replay Specific Messages Without Replaying Everything! 🎯
  112. 112Kafka subscribe() vs assign(): Who Controls Partition Assignment?
  113. 113Kafka Manual Partition Assignment: Using assign() for Control
  114. 114Kafka Rebalance: Avoiding Work Loss or Duplication
  115. 115Kafka Consumer Crash Recovery: Resuming Processing Explained
  116. 116Kafka Consumer State: Position, Offset, and Business State
  117. 117Kafka Stateful Stream Processing: How Does It Remember?
  118. 118Kafka Streams vs Consumer: Choosing the Right Tool
  119. 119Kafka Streams Local State: Fast Processing and Recovery
  120. 120Kafka Streams Windowing: Tumbling vs Hopping Windows Explained
  121. 121Distributed Locks: Preventing Duplicate Work Across Servers
  122. 122Choosing a Distributed Lock: Redis vs Database vs ZooKeeper
  123. 123Leader Election in Distributed Systems Explained
  124. 124Leader Election Algorithms: Bully vs Raft vs ZooKeeper
  125. 125Understanding Consensus in Distributed Systems
  126. 126Understanding Raft Consensus: Terms, Elections & Log Replication
  127. 127Understanding Raft Log Replication: matchIndex vs commitIndex
  128. 128Raft Leader Failure and Re-election Process Explained
  129. 129FAANG HLD 🔥 | Raft Safety — Why Committed Entries Survive Leader Failure! 🛡️
  130. 130Understanding Quorum and Majority in Distributed Systems
  131. 131Paxos Consensus: Understanding Distributed Agreement
  132. 132FAANG HLD 🔥 | Raft vs Paxos — What's the Difference? Consensus Explained! ⚡
  133. 133ZooKeeper Architecture and ZAB Explained
  134. 134Understanding ZooKeeper Ephemeral and Sequential Nodes
  135. 135ZooKeeper Watches: Avoiding the Thundering Herd Problem
  136. 136HLD: Crash Failures vs Network Failures in Distributed Systems
  137. 137HLD: Fail-Stop vs Fail-Recover Explained
  138. 138Understanding Network Partitions in Distributed Systems
  139. 139Understanding Partial Failure in Distributed Systems
  140. 140HLD: Understanding Server Failure Detection and Timeouts
  141. 141Understanding Distributed Clocks in Systems Design
  142. 142Understanding Physical vs Logical Time in Distributed Systems
  143. 143Understanding Lamport Timestamps in Distributed Systems
  144. 144Understanding Vector Clocks in Distributed Systems
  145. 145Understanding Causality in Distributed Systems
  146. 146Understanding RPC: Remote Procedure Calls in Distributed Systems
  147. 147HLD: Understanding REST, RPC, and gRPC Differences
  148. 148Understanding gRPC Architecture: Behind the Call
  149. 149Protobuf and Schema Evolution: Avoid Breaking Changes
  150. 150Understanding Unary vs Streaming RPC in gRPC
  151. 151HLD: Synchronous vs Asynchronous Communication Explained
  152. 152HLD: Understanding Request-Response vs Events in System Design
  153. 153HLD: 5 Microservices Communication Patterns Explained
  154. 154Connection Pooling Explained: Why a Bigger Pool Can Hurt
  155. 155Keep-Alive in Networking: Efficient HTTP Connections
  156. 156When to Use Microservices in Software Design
  157. 157HLD: Microservices vs Modular Monolith Explained
  158. 158Choosing Service Boundaries in Microservices Architecture
  159. 159Understanding Database per Service in Microservices Architecture
  160. 160HLD: Shared DB vs Database per Service Trade-Offs
  161. 161HLD: Circuit Breaker Explained to Prevent Failures
  162. 162Understanding the Retry Pattern in System Design
  163. 163HLD Timeout Pattern: Understanding Timeouts in Distributed Systems
  164. 164Understanding the Bulkhead Pattern in System Design
  165. 165HLD: Fallbacks Explained for System Design
  166. 166Understanding Retry Storms in Distributed Systems
  167. 167Exponential Backoff: Managing Retry Storms in Systems
  168. 168Understanding Jitter in High-Level Design
  169. 169Understanding Circuit Breaker: CLOSED, OPEN, and HALF-OPEN States
  170. 170Cascading Failures in Distributed Systems Explained
  171. 171Load Shedding in Distributed Systems: Protecting Capacity
  172. 172Graceful Degradation in System Design
  173. 173Distributed Backpressure: Managing Service Overload
  174. 174HLD: Fail Fast vs Fail Safe in System Design
  175. 175HLD: Managing Dependency Failures in Distributed Systems
  176. 176Understanding Logs, Metrics, and Traces in Software Engineering
  177. 177Observability in Distributed Systems: Debugging Production Issues
  178. 178Structured Logging: Debugging Production Issues Faster
  179. 179Understanding Correlation IDs in Microservices Architecture
  180. 180Understanding Distributed Tracing in Microservices Architecture
  181. 181Understanding the RED Method for Microservices Monitoring
  182. 182Understanding the USE Method for Infrastructure Bottlenecks
  183. 183Understanding SLIs, SLOs, and SLAs in Reliability Engineering
  184. 184Understanding Error Budgets in High-Level Design
  185. 185Understanding Alerting in High-Level Design
  186. 186HLD: Search Engine Architecture and Its Components
  187. 187Inverted Index: How Search Engines Find Pages Efficiently
  188. 188HLD: Tokenization and Analyzers in Search Engines
  189. 189HLD: Stemming and Normalization in Search Engines
  190. 190HLD: Full-Text Search vs Database Search with Elasticsearch
  191. 191Elasticsearch Architecture: How Distributed Search Works
  192. 192Elasticsearch Shards vs Replicas: Partitioning vs Duplication
  193. 193Elasticsearch Indexing: Why Can't You Search Your Document Yet?
  194. 194Elasticsearch Query Execution: How Searches Work Across Shards
  195. 195Elasticsearch Refresh: Why Your Document Isn’t Searchable Yet
  196. 196Understanding Elasticsearch Relevance Scoring with BM25
  197. 197Elasticsearch Filters vs Queries: MUST vs FILTER Explained
  198. 198Elasticsearch Aggregations: Metrics vs Buckets Explained
  199. 199Elasticsearch Pagination: From & Size vs Search After
  200. 200Understanding Elasticsearch ILM: Hot, Warm, Cold Phases
  201. 201Elasticsearch Aliases and Rollover: Zero-Downtime Index Switching
  202. 202Understanding Elasticsearch Data Streams in HLD
  203. 203Understanding Elasticsearch Index Templates and Components
  204. 204Elasticsearch Mapping: Dynamic vs Explicit Explained
  205. 205Understanding Elasticsearch Shards and Replicas in HLD
  206. 206Elasticsearch Cluster Sizing: How Many Nodes Do You Need?
  207. 207Understanding Elasticsearch Node Roles for Scaling Clusters
  208. 208Elasticsearch Split Brain: Master Election and Quorum Explained
  209. 209Understanding Elasticsearch Snapshots: Replicas vs Backups
  210. 210Understanding Elasticsearch Cross-Cluster Replication (CCR)
  211. 211Understanding Elasticsearch ILM for Cost Management
  212. 212Elasticsearch Data Streams vs. Time-Based Indices: Which to Choose?
  213. 213Elasticsearch Zero-Downtime Reindexing: The Alias Switch Trick
  214. 214Elasticsearch Reindexing: Understanding the Dual-Write Trap
  215. 215Elasticsearch Query Optimization: Filters and Caching
  216. 216Elasticsearch Pagination: Fixing Slow Page Searches
  217. 217Elasticsearch Aggregations: Buckets, Metrics, and Cardinality
  218. 218Elasticsearch Autocomplete: Completion vs. Edge N-Grams vs. Search-as-You-Type
  219. 219Elasticsearch Fuzzy Search: Handling Typos in Queries
  220. 220Understanding Elasticsearch Synonyms and Analyzers
  221. 221Elasticsearch BM25 & Field Boosting: Ranking Search Results
  222. 222Elasticsearch Function Score: Ranking with Business Signals
  223. 223Elasticsearch Rescoring: Improve Search Result Rankings
  224. 224Understanding Elasticsearch Hybrid Search: BM25 vs Vector Search
  225. 225Understanding Elasticsearch Semantic Search and Embeddings
  226. 226Dynamo-Style Databases: High-Level Design Overview
  227. 227Consistent Hashing: Efficient Node Management in Distributed Systems
  228. 228Quorum-Based Replication: How Many Replicas Must Agree?
  229. 229Hinted Handoff: Handling Replica Failures in Databases
  230. 230Cassandra Architecture: Masterless Database Scaling Explained
  231. 231Understanding Cassandra Partition Keys for Data Storage
  232. 232Cassandra Clustering Keys: How Rows Are Sorted
  233. 233MongoDB Architecture: Understanding Replica Sets and Sharding
  234. 234MongoDB Sharding: Scaling to Millions of Orders
  235. 235Time-Series Database Architecture: Storing Millions of Metrics
  236. 236Time-Series Data Modeling: Avoiding Cardinality Issues
  237. 237Understanding Retention Policies in Databases
  238. 238Time-Based Partitioning: Scaling Time-Series Databases
  239. 239HLD: Downsampling and Aggregation in Time-Series Databases
  240. 240Understanding OLTP and OLAP in Database Design
  241. 241Data Warehouse Architecture: Analyzing Business Data Efficiently
  242. 242Data Lake Architecture: Storing Big Data at Scale
  243. 243Data Lakehouse Architecture: Combining Lakes and Warehouses
  244. 244HLD: Batch vs Stream Processing in System Design
  245. 245Stream Processing Architecture: From Events to Real-Time Insights
  246. 246Understanding Watermarks in Stream Processing
  247. 247HLD: Late-Arriving Events - Drop, Update, or Replay?
  248. 248Tumbling vs Hopping Windows in Stream Processing
  249. 249ETL vs ELT: Where Should Data Transformation Happen?
  250. 250Change Data Capture (CDC): Real-Time Database Syncing
  251. 251HLD: Database CDC with Kafka for Real-Time Events
  252. 252Event Sourcing: Store What Happened, Rebuild the State
  253. 253CQRS Explained: Commands vs Queries in System Design
  254. 254HLD: Authentication vs Authorization Explained
  255. 255Session-Based Authentication: How Cookies & Sessions Work
  256. 256Understanding JWT Authentication: How JSON Web Tokens Work
  257. 257OAuth 2.0: Access Data Without Sharing Your Password
  258. 258Understanding OpenID Connect (OIDC) for Authentication
  259. 259Understanding API Keys: Identification and Security
  260. 260Access Tokens vs Refresh Tokens: How Token Renewal Works
  261. 261Understanding Token Expiration and Refresh Token Rotation
  262. 262Understanding mTLS: Microservices Authentication Explained
  263. 263Zero Trust Architecture: Never Trust, Always Verify
  264. 264Encryption at Rest vs In Transit: Protecting Your Data
  265. 265Hashing vs Encryption: Key Differences Explained
  266. 266HLD: Password Hashing and Salt Explained
  267. 267Understanding Cross-Site Scripting (XSS) in Web Security
  268. 268Understanding CSRF: Protecting Against Unwanted Requests

Connected concepts