Troubleshoot the database

Use this guide to identify and repair problems with the Telegraf Controller database. Database errors appear in the Telegraf Controller server log output.

SQLite

Identify the failure type

The following SQLite failures appear in log output, and they require different responses:

  • database is locked: lock contention, not damage. Another connection held the database longer than expected. This is usually transient and resolves on its own. If it persists, verify that only one Telegraf Controller instance uses the database file and that the file is on a local filesystem. Do not run repair commands for lock errors.
  • database or disk is full: the volume holding the database file is out of space, not damaged. Free disk space and restart Telegraf Controller. No repair is needed unless a corruption error also appears.
  • unable to open database file: the server cannot read or create the database file, which is usually a path or permissions problem, not damage. Check that the database directory exists and that the user running Telegraf Controller can read and write the file and its directory. Do not run repair commands for this error.
  • database disk image is malformed: database corruption. The database file or one of its internal structures is damaged. Corruption does not heal on its own and the affected queries keep failing until you repair the database. Follow the rest of this section.

Prerequisites: install the SQLite CLI

Diagnosis and repair use the sqlite3 command-line shell, which operates directly on the database file:

  • macOS: included with the operating system.
  • Linux: install the sqlite3 package, for example apt install sqlite3 or dnf install sqlite.
  • Windows: download the sqlite-tools bundle from the SQLite download page.

Stop Telegraf Controller before running any diagnosis or repair command. Repairing a database while the server writes to it can make the damage worse.

Diagnose corruption

  1. Stop Telegraf Controller.

  2. Run an integrity check against the database file. Replace /path/to/sqlite.db with your database location (for default locations, see Default SQLite data locations):

    sqlite3 /path/to/sqlite.db "PRAGMA integrity_check;"
  3. Interpret the output:

    • ok: the database is intact. The error came from something else, for example a permissions problem or a full disk.

    • Index errors only, for example:

      wrong # of entries in index some_index_name

      Only index structures are damaged and the underlying table data is intact. This is repairable in place with no data loss. Continue to Repair corrupted indexes.

    • Other errors, for example messages that name table pages or rows: table data itself is damaged. Continue to Recover from severe corruption.

Repair corrupted indexes

If the integrity check reported only index errors, rebuild all indexes from the intact table data:

  1. Back up the damaged database file before changing it.

  2. Rebuild the indexes:

    sqlite3 /path/to/sqlite.db "REINDEX;"
  3. Verify the repair:

    sqlite3 /path/to/sqlite.db "PRAGMA integrity_check;"

    The check should now return ok.

  4. Start Telegraf Controller and confirm the log output no longer reports database errors.

Recover from severe corruption

If the integrity check reports damage beyond indexes, recover what SQLite can read into a new database file:

  1. Back up the damaged database file.

  2. Run the recovery:

    sqlite3 /path/to/sqlite.db ".recover" | sqlite3 /path/to/recovered.db
  3. Check the recovered file:

    sqlite3 /path/to/recovered.db "PRAGMA integrity_check;"
  4. Replace the damaged database with the recovered file, keeping the original for reference:

    mv /path/to/sqlite.db /path/to/sqlite.db.damaged
    mv /path/to/recovered.db /path/to/sqlite.db
  5. Start Telegraf Controller and verify your configurations, agents, and users in the web interface.

Recovery salvages everything SQLite can still read, but rows in damaged regions may be lost. If recovery produces incomplete data, restore from a backup instead.

Prevent corruption

SQLite is robust against crashes in normal operation. Corruption almost always traces back to one of the following, all avoidable:

  • Copying a live database: never copy the database file while Telegraf Controller runs. Use the safe backup methods instead.
  • Deleting companion files: never delete or move the -wal and -shm files next to the database. They are part of the database.
  • Network filesystems: keep the database file on a local filesystem. File locking on NFS and SMB shares is unreliable and can corrupt SQLite databases.
  • Forced shutdown: stop Telegraf Controller with a normal termination signal (SIGTERM or Ctrl+C) rather than kill -9, and avoid powering off the host while the server is writing.
  • No backups: If you can’t repair database corruption and don’t have a backup, you lose data. [Back up the database](/telegraf/controller/admin/database/back-up-and-restore/) on a schedule.

PostgreSQL

Telegraf Controller relies on your PostgreSQL server for database health, so troubleshooting is directed at the server rather than at Telegraf Controller:

  • Connection problems: verify that the PostgreSQL server is running, check the format of and credentials in your connection string (DSN or database URL), and verify network connectivity between the Telegraf Controller host and the server.
  • TLS handshake failures: error performing TLS handshake in the log means Telegraf Controller does not trust the certificate presented by the PostgreSQL server. Provide the certificate authority (CA) certificate that signed it. See Provide the database CA certificate.
  • Server health and corruption: use your PostgreSQL tooling and the PostgreSQL documentation.
  • Unrecoverable state: restore from a backup.

Was this page helpful?

Thank you for your feedback!