تقع البيانات الزمنية في صميم العديد تطبيقات إنترنت الأشياء. Devices generate readings in near real-time, producing large volumes of timestamped data that must be stored efficiently, ingested quickly, and retrieved easily for analysis. While traditional relational databases can store time-series data, they are not always optimised for the unique access patterns required – such as time-based aggregations or rapid ingestion with minimal delays. So, in this article, I will focus on choosing the best time-series database for specific needs.

أدناه، سأقارن أربعة حلول شائعة لتخزين وتحليلات السلاسل الزمنية:

  • TimescaleDB – امتداد فوق PostgreSQL، يوفر تحسينات للسلاسل الزمنية مع الحفاظ على ألفة SQL.
  • InfluxDB (v3) – قاعدة بيانات سلاسل زمنية مصممة خصيصًا ومعروفة بقدرتها العالية على الاستيعاب ونظامها البيئي الغني.
  • Azure Data Explorer (ADX) – خدمة تحليلات سريعة أصلية سحابيًا من Microsoft Azure، مُحسّنة لبيانات السجلات والقياس عن بُعد.
  • AWS Timestream – خدمة قاعدة بيانات سلاسل زمنية مُدارة بالكامل من AWS، مصممة للتوسع بأقل قدر من الأعباء التشغيلية.

على الرغم من أن TimescaleDB وInfluxDB تقدم أيضاً عروضاً سحابية، إلا أنني سأركز في هذا المقال على توفرهما كحلول سلاسل زمنية محلية.

At Spyrosoft, we always consider these options at the beginning of a new IoT project, based on the specific requirements. The goal of this article is to compare these solutions for IoT data storage and analytics. To be clear, there is no single “golden rule” or universal best choice when selecting a time-series database. However, I will present the key considerations you should keep in mind when choosing a time-series solution, by providing a comparative overview.

لنبدأ!

تنظيم البيانات في قاعدة البيانات الزمنية

يُعدّ التنظيم الفعّال للبيانات أمرًا أساسيًا لأداء أفضل قواعد البيانات الزمنية وتحسينها. وتُعدّ حلول مثل InfluxDB, TimescaleDB, Amazon Timestream، و Azure Data Explorer يمكن لكل منها استخدام مخططات مميزة لهيكلة وتخزين بيانات السلاسل الزمنية بفعالية.

InfluxDB Clustered

يتيح لك InfluxDB تخزين البيانات في موقع مُسمّى يُطلق عليهقاعدة البيانات (يُشار إليه بـالفئات في InfluxDB TSM)، الذي يجمع البيانات منطقيًا في الجداول (المعروف باسمالقياسات في InfluxDB TSM). يحتوي كل جدول على الوسوم و الحقول:

  • الوسوم هي أزواج مفتاح-قيمة توفر بيانات وصفية لكل نقطة – تشمل الأمثلة معرّفات مثل المحطة أو معرّف المستشعر أو الموقع. قد تكون قيم الوسوم فارغة.
  • الحقول are key-value pairs representing values that change over time – examples include temperature or pressure. Field values may be null, but at least one field value must be non-null in any given row.

A الطابع الزمني (which is never null) is associated with each data point, and all data is ordered by time. The term “point” refers to a single data record identified by its measurement, tag keys, tag values, field key, and timestamp. All points in a given table should share the same tags. The columns that uniquely identify each row in a table form the المفتاح الأساسي. يتم تحديد الصفوف بشكل فريد من خلال الطابع الزمني ومجموعة الوسوم غير الفارغة.

عند كتابة البيانات إلى InfluxDB، تُعرّف البيانات نفسها المخطط. لا حاجة لإنشاء جداول أو تعريف مخطط مسبقاً بشكل صريح.

TimescaleDB

تم بناء TimescaleDB على PostgreSQL ويتم توزيعه كإضافة لـ PostgreSQL، مع الحفاظ على دعم SQL الكامل. يقوم الحل بتنظيم بيانات السلاسل الزمنية فيالجداول الفائقة، وهي في جوهرها جداول PostgreSQL مقسّمة حسب الوقت. وتدير قاعدة البيانات هذه الأقسام تلقائياً خلف الكواليس. ويتكون الجدول الفائق (hypertable) من جداول أصغر تُسمى الكتل، يُخصص لكل منها نطاق زمني لتخزين البيانات من تلك الفترة فقط. ال حجم الكتلة configures itself during the hypertable creation, so it should be carefully planned, as it affects insert and query performance. By default, a newly created hypertable indexes by time in descending order. Hypertables can coexist with standard PostgreSQL tables, which can be advantageous in certain scenarios.

Timestream

يخزن Amazon Timestream البيانات فيقواعد البيانات التي تحتوي على الجداول، على غرار بنية InfluxDB. يحتوي كل جدول علىالسلاسل الزمنية، وهي عبارة عن تسلسل من نقطة بيانات واحدة أو أكثر (سجلات) يتم التقاطها خلال فترة زمنية. وتُسمى نقطة البيانات الواحدة في السلسلة الزمنيةالسجل.

السمة التي تصف البيانات الوصفية لسلسلة زمنية تُعرف باسمالبُعد، وهو يتكون من اسم وقيمة (على سبيل المثال، "device_id" و"12345"). أ قياس هي قيمة يتتبعها السجل، تُحدد باسم القياس وقيمة القياس (على سبيل المثال، "درجة الحرارة" و"45"). أما الطابع الزمني يشير إلى وقت جمع القياس، بدقة تصل إلى النانوثانية.

Azure Data Explorer

الحاوية العلوية في Azure Data Explorer هي قاعدة بيانات، تحتوي على الجداول. يخزّن كل جدول البيانات فيالنطاقات (data shards). An extent is a table’s horizontal segment containing data and metadata, such as its creation time and optional tags. All extents together form the table. They are also evenly distributed across cluster nodes and cached in both local SSDs and memory for optimal performance. Essentially, they are immutable, and each extent physically stores records in columns.

الاستعلام عن البيانات باستخدام أفضل قواعد بيانات السلاسل الزمنية

A table presenting a comparison of querying data for choosing the best time-series database.

طرق الإدخال لقواعد البيانات الزمنية

تتنوع استراتيجيات الإدخال (Ingestion) بشكل كبير عبر قواعد البيانات الزمنية هذه، مما يعكس الاحتياجات المتنوعة لتطبيقات إنترنت الأشياء (IoT).

  • InfluxDB supports high-throughput writes through its line protocol via HTTP, as well as integrations with tools like Telegraf (a server-based agent that can collect and send metrics and events from IoT sensors) for streaming and batch imports.
  • TimescaleDB، بكونه امتداداً لـ PostgreSQL، يستفيد من عمليات إدراج SQL القياسية، وعمليات النسخ المتوازي المجمعة timescaledb-parallel-copy لاستيراد البيانات، على سبيل المثال من ملفات CSV، والموصلات الخارجية.
  • AWS Timestream يوفر عمليات تكامل أصلية مع AWS IoT Core وKinesis Data Streams، مع تقديم نهج قائم على SDK في الوقت نفسه.
  • Azure Data Explorer (ADX) يمكنه استيعاب البيانات من Event Hubs أو IoT Hub أو نقاط النهاية المباشرة القائمة على HTTP، مع تجميع البيانات وإدارة أجزائها تلقائيًا.

أفضل قاعدة بيانات سلاسل زمنية: خيارات الاستضافة

تقدم كل قاعدة من قواعد البيانات الزمنية هذه خيارات استضافة مختلفة، مما قد يؤثر على التكلفة وقابلية التوسع والتعقيد التشغيلي.

InfluxDB

There are many ways to implement InfluxDB into IoT solutions: as a self-hosted on-premises instance, in a private cloud, or using InfluxDB Cloud, a fully managed SaaS offering. The self-hosted version provides complete control over infrastructure but requires operational management. For this article, we focus on InfluxDB hosted on-premises. You can deploy a single instance of InfluxDB, or utilise InfluxDB Clustered, designed with high availability and scalability in mind. Deploying InfluxDB Clustered on Kubernetes requires additional resources, such as persistent storage for underlying Parquet files that must be compatible with AWS S3 or S3-compatible object storage, and an external PostgreSQL (or PostgreSQL compatible) instance for metadata and coordination. It is also advisable to use a load balancer to efficiently distribute queries and ingestion requests across cluster nodes.

TimescaleDB

TimescaleDB is available as a self-managed PostgreSQL extension, making it deployable on any infrastructure where PostgreSQL runs. Additionally, TimescaleDB offers Timescale cloud, a managed service for hosting and scaling time-series databases with minimal operational overhead.

AWS Timestream

AWS Timestream is a fully managed, cloud-native service that is exclusively available within the AWS ecosystem. It eliminates the need for infrastructure management but requires AWS integration and follows a cloud-based pricing model. At the time of writing this article, AWS Timestream employs a pay-as-you-go pricing structure, with costs determined by data ingestion, storage, and query processing. For the most up-to-date pricing details, you should refer to the official AWS Timestream pricing page. Pricing should be assessed based on specific project requirements, as costs may vary depending on usage patterns and data retention needs.

Azure Data Explorer

ADX is a cloud-native service that runs on Microsoft Azure, offering a managed environment with built-in scalability. While it is optimised for Azure workloads, it can also integrate with hybrid and multi-cloud architectures through various ingestion methods. Azure Data Explorer provides multiple service tiers, including a Dev/Test cluster, which is designed for development and testing with a single node and no redundancy, and a Production cluster, which includes at least two nodes for high availability and operates under an Azure Data Explorer SLA. You should select an appropriate tier based on your workload requirements and cost considerations.

أفضل قاعدة بيانات سلاسل زمنية واحدة هي… أم أنها ليست كذلك؟

حسناً، الأمر يعتمد على الحالة! لا يوجد خيار واحد مثالي يناسب جميع حالات الاستخدام. ومع ذلك، إليك دليلاً تقريبياً يوضح متى قد تكون كل قاعدة بيانات هي الخيار الأفضل:

  • إذا كنت بحاجة إلى استضافة محلية – فكّر في InfluxDB أو TimescaleDB.
  • إذا كان فريقك على دراية بـ SQL ويفضل PostgreSQL – فقد تكون TimescaleDB هي الخيار الأمثل لك.
  • إذا كنت بحاجة إلى حل مُدار بالكامل وقائم على السحابة على AWS – فإن AWS Timestream خيار طبيعي.
  • إذا كانت بنيتك التحتية قائمة على Azure وكنت بحاجة إلى تكامل سلس مع موارد Azure الأخرى، مثل Data Lake وPower BI – فإن Azure Data Explorer يستحق النظر فيه.
  • إذا كانت معدلات الاستيعاب العالية والتحليلات في الوقت الفعلي أمرًا بالغ الأهمية – فقد يكون InfluxDB أو Azure Data Explorer خيارك الأفضل.

Of course, these are just starting points, and the best approach is always to conduct a proof of concept (PoC) and load tests tailored to your specific project. After all, picking a database is a bit like picking a favourite pizza topping – what works for one team may not be the best for another.

اكتشف كيف يمكننا الارتقاء بحلول إنترنت الأشياء الخاصة بك

اعرف المزيد

استنتاجات حول اختيار أفضل قاعدة بيانات سلاسل زمنية

There is no single “golden rule” for choosing the best time-series database for an IoT project. Each presented solution has its strengths and trade-offs, and the best option depends on the project’s specific requirements, such as scalability, ease of querying, and data acquisition performance. Nevertheless, you should definitely consider the factors outlined in this post when making a decision.

The good news is that deploying these databases is relatively straightforward, making it easy to conduct a proof of concept (PoC) that evaluates their performance in a real-world scenario. Once a PoC is in place, artificial load testing can help estimate how well a database handles expected workloads, ensuring the right choice before committing to a production system.

In my current project, we have followed this approach from the beginning. When we gathered telemetry requirements from the customer, we initially conducted a PoC with both Azure Data Explorer (ADX) and PostgreSQL. Given that the customer already had an existing Azure infrastructure and additional requirements – such as integration with other resources like Data Lake and Power BI – ADX emerged as the best fit for our scenario. It provided the necessary scalability almost immediately, with seamless integration into services like IoT Hub and Blob Storage.

Another important consideration was data migration from different sources. The ability to include external tables – such as sources stored in CSV files within Blob Storage – was also a significant advantage for our project’s needs. However, this does not mean that Azure Data Explorer is the best choice for every project. It was simply the most suitable option for us, considering our specific requirements.

Of course, it wasn’t all smooth sailing from the start. We had to refine our ingestion approach, configure data aggregation, manage data latency, and carefully consider data retention and continuous export strategies. Additionally, cost efficiency was a bit tricky – pricing for ADX isn’t always straightforward and requires careful analysis. But handling ADX and making it work efficiently is probably a topic worthy of its own post (and maybe even a few deep sighs in the process).

استفد من أفضل قاعدة بيانات سلاسل زمنية مع Spyrosoft

Choosing the best time-series database for your IoT project is no easy task, but hopefully, this comparison has given you some valuable insights to guide your decision. Whether you’re looking for seamless cloud integration, SQL familiarity, or high-throughput ingestion, each database has something unique to offer.

إذا كنت بحاجة إلى مزيد من الخبرة في إنترنت الأشياء أودعم عملي في تطوير مشروع إنترنت الأشياء الخاص بك، تواصل معنا عبر النموذج أدناه، واكتشف ما يمكننا تحقيقه معاً.

الأسئلة الشائعة

A time-series database is designed to store data points collected over time, often at high frequency. Unlike relational databases, which focus on structured, relational datasets, time-series systems optimise for fast ingestion, efficient compression, and smooth querying of sequential data. This makes them much better suited for IoT scenarios where devices generate continuous streams of measurements.

Focus on ingestion speed, storage efficiency, query performance, and how well the database scales as data volumes grow. It also helps to consider ease of integration with your existing infrastructure and whether the system supports features such as downsampling or retention policies. The best time-series database for your project will match both your current load and your future growth.

IoT environments rarely stay static. Device numbers increase, sampling frequencies change, and new use cases appear. A database that scales without disruption allows you to maintain consistent performance and predictable costs as your environment expands. Without this, even well-designed systems can become slow or unstable.

Many open-source solutions provide strong foundations, large communities, and stable features. Their transparency also makes them easy to audit and customise. However, industrial projects may need additional safeguards, such as defined SLAs, enterprise support, or certified integrations. It often comes down to your internal capabilities and long-term maintenance plans.

Data retention policies determine how long you keep raw and processed data. A suitable policy helps you control storage costs without losing valuable insights. Some databases automate retention and downsampling, which is useful when handling long-term trends without storing unnecessary detail.

Slow queries can reduce the value of real-time monitoring. A database built for time-series workloads will handle aggregations, filtering, and window functions more smoothly. This gives engineers faster access to insights, supports alerting, and helps detect anomalies before they become costly issues.

Yes. Many organisations use a hybrid approach. For example, a time-series database might store high-frequency sensor data, while a relational or document database manages metadata, reports, or business logic. The important part is to design data flows that remain stable and clear as the system evolves.

You can rely on vendor resources, community documentation, or external consultancy. According to Spyrosoft, organisations benefit from having a partner who understands both the technical landscape and the practical demands of large-scale IoT platforms. Guidance of this kind can help you choose a solution that fits your long-term strategy.