For the complete documentation index, see llms.txt. This page is also available as Markdown.

List of alerts and corrective actions

Check WEKA system alerts and take necessary actions based on severity and nature.

Explore the list of system alerts, their meanings, and the recommended corrective actions. Understanding these alerts helps you respond to system events effectively.

For detailed maintenance and troubleshooting articles, search the WEKA Knowledge Base in the WEKA support portal. Access to the portal requires a login and is available for registered customers and internal users. You can also contact the Customer Success Team.

The LicenseError alert includes a 5% grace threshold. It is triggered only when used capacity exceeds 105% of the licensed limit, which allows for minor overages without generating an immediate alert.

Alert Name
Description
Corrective Action
Severity

AdminDefaultPassword

Default admin password in use

Change the admin password to restrict access to authorized users.

WARNING

AgentNotRunning

Local agent is not running

Restart the local agent on the specified server using weka agent restart.

DEBUG

ApproachingClientsUnavailability

Approaching connected clients limit

Ensure all backend containers are up. Alternatively, expand the cluster with additional backend containers or servers.

DEBUG

ApproachingSystemLimit

Approaching a system limit

{action_item}

MAJOR

AutoRemoveTimeoutTooLow

Stateless client auto-remove timeout too low

Remount the host with a higher auto-remove timeout value.

WARNING

AvailableMemory

Not enough available memory

Check the system memory usage.

MAJOR

BackendNumaBalancingEnabled

NUMA balancing is enabled on a backend server

Disable automatic NUMA balancing on the backend server: echo 0 > /proc/sys/kernel/numa_balancing.

WARNING

BackendVersionsMismatch

Backend version does not match the cluster version

Upgrade all backends to match the cluster version.

WARNING

BadDisksCapacityRatio

Large discrepancy between the smallest and largest drive

Replace drives with comparable capacities to minimize the size difference between drives.

MAJOR

BlockedJrpcMethod

JRPC method is blocked

Unblock the JRPC method using the blocked_jrpc_methods_remove or blocked_jrpc_methods_clear manhole.

DEBUG

BondInterfaceCompromised

Network high-availability interface compromised

Verify the network configuration, cables, and NICs.

MINOR

BucketCapacityExhausting

Buckets are approaching maximum capacity

Migrate to a cluster with a higher number of buckets.

DEBUG

BucketHasNoQuorum

Too many compute processes are down

Ensure the compute processes on containers {containers} are running and connected. If the issue persists, contact the Customer Success Team.

DEBUG

BucketUnresponsive

Compute resource failure

Check connectivity and drive status on the leader container. Ensure compute processes are running. If the issue persists, contact the Customer Success Team.

CRITICAL

BucketsImbalance

FS buckets are not properly balanced between processes

If this alert persists, contact the Customer Success Team.

WARNING

CPUFrequentStarvation

Frequent CPU starvation detected in the last minute

Check the logs of the relevant containers for potential hardware or core allocation problems.

DEBUG

CPUStarvation

CPU starvation detected in the last minute

Check the logs of the relevant containers for potential hardware problems. For a specific hang address, run /weka/weka_addr2line within the reported WEKA container on the address to convert it into a symbol name.

DEBUG

CWTaskAbortionStuck

CWTask stuck in aborting state

Start I/O to allow the task to complete the abort process.

DEBUG

ChokingDetected

High congestion level

Expand cluster resources if the workload is maximizing current capacity. For assistance, contact the Customer Success Team.

DEBUG

ClientVersionsMismatch

Client version does not match the cluster version

Upgrade clients to match the cluster version using weka local upgrade.

DEBUG

ClockSkew

Clock skew detected on server

Ensure NTP is configured correctly on the containers and that their clocks are synchronized.

MINOR

CloudHealth

WEKA Home disconnected

Check that the server has Internet connectivity and is connected to WEKA Home. See the WEKA Home – The WEKA Support Cloud topic.

MINOR

CloudStatsError

Statistics upload failed

See the event details in System Events.

DEBUG

ClusterInitializationError

Cluster initialization error

Identify and resolve the underlying cause to resume I/O operations. To clear this alert, run weka cluster stop-io.

WARNING

ClusterIsUpgrading

Cluster upgrade in progress

Wait for the upgrade to complete. If the upgrade fails, contact the Customer Success Team.

DEBUG

CompressedSpaceFragmented

Compressed free space is too fragmented

Check whether DATA_DEFRAG tasks are paused or hanging. If the issue persists, contact the Customer Success Team.

WARNING

ConfigOverridesActive

Configuration overrides are active

Contact the Customer Success Team.

DEBUG

CoreOverlapping

Core overlapping detected

Contact the Customer Success Team.

MAJOR

DataIntegrity

Data integrity issues detected

Contact the Customer Success Team with scan details. Avoid manual repairs without guidance to prevent further data loss.

CRITICAL

DataProtection

Partial data protection

Restore full data protection. If the cluster remains resilient to one failure, identify the down process, container, or drive and resolve the issue.

MINOR

DataServiceConfigurationNotSet

Dataservice global configuration filesystem is not set

Set config_fs for the dataservice by running weka dataservice global-config set --config_fs.

MINOR

DedicatedWatchdog

Hardware watchdog driver required for dedicated server

Install the driver by following the Knowledge Base instructions in the WEKA Support portal.

DEBUG

DrainingStuck

Host stuck in draining state

Check the host status and logs to identify blockers such as slow data migration, capacity constraints, or I/O errors. Verify cluster health and available capacity. If the issue persists, contact the Customer Success Team.

MINOR

DriveAutoEjected

Drive automatically ejected

Monitor the drive phase-out progress and consider planning a drive replacement.

WARNING

DriveAutoReactivating

Drive auto-reactivation in progress

Monitor the drive status. {reactivation_action}

WARNING

DriveCriticalWarnings

Drive critical warnings detected

Deactivate the drive using weka cluster drive deactivate and replace it.

MAJOR

DriveDown

Drive failed

Contact the Customer Success Team to assess drive health and determine whether physical replacement is required. Provide the drive UUID, error message, and host details for faster diagnosis.

MINOR

DriveEndurancePercentageUsed

Drive has exceeded its life expectancy

Replace the specified drive before it fails.

MAJOR

DriveEnduranceSparesRemaining

Drive internal spares are running low

Replace the specified drive before it fails.

MAJOR

DriveNVKVRunningLow

Drive is nearing exhaustion of an internal resource

Contact the Customer Success Team.

DEBUG

DriveNeedsPhaseout

Drive has too many errors

Deactivate the drive using weka cluster drive deactivate {deactivate_flag} and consider replacing it.

MAJOR

ExampleAlert

Example alert

Disable this alert by running the set_example_alert_off manhole.

DEBUG

ExceptionsDuringAlertsEvaluation

Exceptions thrown during alerts evaluation

Check the Assertion Failures event to identify the source of the problem. If the issue persists, contact the Customer Success Team.

DEBUG

FaultsEnabled

Faults are enabled

Contact the Customer Success Team.

DEBUG

FilesystemKMSError

Filesystem KMS error

Review the filesystem's KMS customization, configuration, and connectivity.

DEBUG

FilesystemsThinProvisioningLowSpace

Filesystems with thin provisioning are running low on space

Consider adding more SSD capacity to the tenant containing these filesystems.

WARNING

FilesystemsThinProvisioningReserveReached

Filesystems thin provisioning capacity reserve has been reached

Create a new filesystem or expand the filesystem's capacity using the reserved capacity.

DEBUG

GroupQuotasHardLimitReached

Group quota hard limit exceeded

Run weka fs quota list --group to list groups that have exceeded their hard quota limits. Free up space for these groups or increase their hard quota limit.

WARNING

GroupQuotasSoftLimitReached

Group quota soft limit exceeded

Run weka fs quota list --group to list groups that have exceeded their soft quota limits. Free up space for these groups or increase their hard quota limit.

WARNING

HangingCacheSync

Cache sync is hanging

Consider using weka debug fs drop-dirty-cache to drop the cache and allow other clients to access the file. Note that unsynchronized writes will be lost.

MINOR

HangingClusterTasks

Cluster background task progress is hanging

Contact the Customer Success Team to address the stalled task.

DEBUG

HangingIos

Some I/Os have stopped responding

Ensure the compute processes are up, running, and connected. If a backend object store is configured, ensure it is connected and responsive. If the issue persists, contact the Customer Success Team.

DEBUG

HardwareComponentFaulty

Hardware component fault detected

Check the management controller (BMC or IPMI) for diagnostic details and take corrective action.

MAJOR

HardwareComponentWarning

Hardware component warning detected

Check the management controller (BMC or IPMI) for details on the affected components.

WARNING

HardwareMonitoringUnavailable

Hardware monitor unavailable

Verify that BMC or IPMI settings are configured correctly in container resources. If the issue persists, contact the Customer Success Team.

MINOR

HighDrivesCapacity

SSD capacity overflow

Free up space on the SSDs or add more SSDs to the cluster. See the Expand Specific Resources of a Container topic.

MAJOR

HighLevelOfUnreclaimedCapacityInObjectStore

High level of unreclaimed space in an object store

Check object store connectivity and the progress of deletion operations. Verify that deletion operations are authorized on the object store. Run weka fs tier capacity --filesystem {filesystem} for details.

DEBUG

HighSSDToRAMRatio

High SSD-to-RAM ratio

Consider increasing RAM cluster-wide or removing unneeded drives to meet filesystem (RAID) requirements and lower the SSD-to-RAM ratio.

DEBUG

HotspotInodes

Some files have a long I/O waiting queue

Retrieve the file path using weka debug fs resolve-inode {inodeIdValue} --snap-view-id {snapViewIdValue}. List clients with open handles using weka debug fs lsof {inodeIdValue} --snap-view-id {snapViewIdValue}. Use the access pattern to optimize the application's caching mechanism.

DEBUG

IBNotEnhanced

Enhanced IB mode is disabled

Contact the Customer Success Team.

DEBUG

ImbalancedCpuUsage

Imbalanced CPU usage detected in cluster processes

Examine the system configuration for abnormalities that may be causing the CPU usage imbalance.

DEBUG

JumboConnectivity

A container cannot send jumbo frames

Check the container network settings and the connected switch, and enable jumbo frames to improve performance.

WARNING

KMSError

KMS error

Review the KMS configuration and connectivity.

MAJOR

LeaderPreparedForUpgrade

Leader prepared for upgrade

After the upgrade, the leader state returns to normal automatically. If the issue persists, contact the Customer Success Team.

DEBUG

LegacyManualOverridesActive

Legacy manual overrides are active

Contact the Customer Success Team.

DEBUG

LicenseError

License error

Ensure the cluster uses the correct license, that the license has not expired, and that the allocated space does not exceed the license limits.

WARNING

LocalTLSCertificateExpired

Local TLS certificate has expired

Update the local certificate.

DEBUG

LocalTLSCertificateExpiringSoon

Local TLS certificate is expiring soon

Update the local certificate.

DEBUG

LocalTLSConnectivityToNeighbors

Outgoing TLS connectivity to backends is down

Resolve the TLS issue. Check for errors in the local CA certificate at /etc/wekaio/certs.

DEBUG

LongestWaitInodes

Elevated I/O wait time detected for some files

Retrieve the file path using weka debug fs resolve-inode {inodeIdValue} --snap-view-id {snapViewIdValue}. List clients with open handles using weka debug fs lsof {inodeIdValue} --snap-view-id {snapViewIdValue}. Use the access pattern to optimize the application's caching mechanism.

DEBUG

LowDiskSpace

Low disk space

See the event details in System Events.

MINOR

ManualOverridesActive

Manual overrides are active

Contact the Customer Success Team.

DEBUG

ManualOverridesForced

Manual overrides are forced

Contact the Customer Success Team.

DEBUG

MismatchedDriveFailureDomain

A drive's failure domain does not match that of its attached container

Attach the drive to a container with a matching failure domain, or re-provision the drive to clear its failure-domain assignment.

MAJOR

MismatchedJoinSecrets

Backend containers do not have the same join secrets

Ensure all backend containers have the same join secrets to avoid problems rejoining or reforming the cluster.

DEBUG

NegativeUnprovisionedCapacity

Negative unprovisioned capacity

Resize one or more filesystems to reclaim capacity. For more information, contact the Customer Success Team.

DEBUG

NetworkFailedToStartPorts

Network ports failed to start

Run weka debug net ports $NODE to view the current port status.

DEBUG

NetworkInterfaceLinkDown

Network interface link is down

Check connectivity to the specified network interface and verify that it is not blocked.

MINOR

NfsLocksDisabled

NFS locks are disabled

Configure the config filesystem using weka nfs global-config set --config-fs=<filesystem name>.

WARNING

NfsServiceDownAlert

NFS service is down

If the service remains down, contact the Customer Success Team.

MAJOR

NoCgroupsConfigured

No cgroups are configured

Enable and configure cgroups (v1/v2) following the Cgroups configuration section in the documentation.

WARNING

NoClusterLicense

No license assigned

Obtain and install a license from get.weka.io.

WARNING

NoHotSpareFailureDomains

No hot-spare failure domains have been configured

Assign one or more hot-spare failure domains by running weka cluster hot-spare.

WARNING

NodeBlacklisted

A process cannot rejoin the cluster

Whitelist the process by running weka debug blacklist disable to allow it to rejoin the cluster.

DEBUG

NodeDisconnected

Process disconnected

Check network connectivity to ensure the processes can communicate with the cluster.

MINOR

NodeNetworkUnstable

A process with an unstable network connection was detected

Ensure proper network connectivity in the cluster. If the issue persists, contact the Customer Success Team.

WARNING

NodeRDMANotActive

RDMA support for the process is inactive

Ensure that at least one RDMA-capable device exists.

DEBUG

NodeTieringConnectivity

A process cannot connect to an object store

Check object store connectivity and ensure the process can communicate with it. Verify network connectivity, process health, and the object store vendor equipment.

MAJOR

NonTlsApisAllowed

Non-TLS APIs are allowed

Update TLS settings to enforce encrypted API communication and disable non-TLS access over HTTP.

DEBUG

NotEnoughActiveDrives

Reduced data protection

Check connectivity and server status. Activate drives in additional failure domains.

MAJOR

NotEnoughMemoryForFilesystemOperation

Insufficient cluster-wide RAM for proper filesystem operation

Increase RAM cluster-wide to meet filesystem (RAID) requirements, or remove drives to reduce SSD capacity.

WARNING

NotEnoughSSDCapacity

Some provisioned capacity is unavailable due to failed drives

Check for drives that are down.

MAJOR

NotificationQueueHighLoad

S3 notification queue has reached the high-water mark

Review the queue load and address the backlog.

MINOR

NotificationSendFailure

S3 notification send failure

Check system logs for details and restore Kafka service availability.

MAJOR

PartialConnectivityTrackingDisabled

Partial connectivity tracking is disabled

Contact the Customer Success Team to re-enable partial connectivity tracking.

DEBUG

PartialHugepageAllocation

Not enough memory for hugepage allocation

Check the system memory configuration.

MAJOR

PartiallyConnectedNode

A partially connected process was detected

Ensure proper network connectivity in the cluster. If the issue persists, contact the Customer Success Team.

MINOR

PassedClientsAvailabilityThreshold

Connected clients limit reached

Add more backend containers or servers to the cluster, check whether any backends are down, or disconnect some clients.

DEBUG

PathsDegraded

Network paths are degraded

Contact the Customer Success Team to review path connectivity.

MINOR

PerformanceDegradedLowRAM

Low server RAM

Add more servers to the cluster, add RAM to the backend servers, or increase the memory allocation to the compute processes.

MAJOR

QuotasHardLimitReached

Directory quota hard limit exceeded

Run weka fs quota list to list directories that have exceeded their hard quota limits. Free up space for these directories or increase their hard quota limit.

WARNING

QuotasSoftLimitReached

Directory quota soft limit exceeded

Run weka fs quota list to list directories that have exceeded their soft quota limits. Free up space for these directories or increase their hard quota limit.

WARNING

RAIDCapacityExhaustion

RAID capacity nearing exhaustion

Monitor the situation for approximately 5 minutes to allow for automatic rebalancing. If RAID capacity usage remains high or continues to increase, contact the Customer Success Team to prevent write failures.

MAJOR

RequestedActionFailure

Requested action failed

Check system logs for details and verify host accessibility, resource availability, and system state before retrying.

DEBUG

RequestedActionTimeout

Container action has timed out

Check the container status and logs for more information about the timeout.

DEBUG

ResourcesNotApplied

Resource changes not applied

Apply the resource changes by running weka cluster container apply <host-id>.

DEBUG

S3LifecycleTaskCannotStart

S3 lifecycle tasks cannot start due to missing dataservice containers

Set up dataservice containers following the "Set up a Data Services Container for Background Tasks" documentation.

MINOR

S3UpgradeInProgress

S3 is undergoing a software upgrade

If the upgrade does not complete successfully, contact the Customer Success Team.

WARNING

SSDCapacityDiscrepancy

Mismatch between actual and expected SSD capacity usage

Monitor compute process stability. If the mismatch persists, contact the Customer Success Team.

DEBUG

SSDCapacityTooHigh

Available capacity cannot be fully utilized

Contact the Customer Success Team to optimize SSD capacity usage.

WARNING

SystemDefinedTLS

TLS certificate is not user-defined

Replace the auto-generated self-signed certificate with a user-defined certificate by running weka security tls set.

WARNING

SystemdSafeShutdownConfigMismatch

Systemd safe shutdown configuration differs across cluster servers

Enable systemd safe shutdown on all servers by running weka local enable-safe-shutdown.

WARNING

TLSCertificateExpired

TLS certificate has expired

Replace the existing certificate using weka security tls set.

MAJOR

TLSCertificateExpiresSoon

TLS certificate is about to expire

Replace the existing certificate using weka security tls set.

MAJOR

TelemetryStatusFault

Telemetry is not streaming

Check the telemetry sinks configuration.

WARNING

TieredFilesystemOverfillingSSD

Tiered filesystem SSD capacity threshold exceeded

Verify tiering service connectivity to object storage and check for network or authentication issues blocking data migration. Consider increasing tiering upload bandwidth or reviewing tiering policies and thresholds if the SSD is consistently near capacity.

WARNING

TooManyPendingClusterwideJobs

Too many pending cluster-wide jobs

Consider adjusting the policy configuration.

DEBUG

TraceDumperDown

Trace dumper is down

Contact the Customer Success Team to restart the trace dumper.

DEBUG

TracesDisabled

Traces are disabled

Re-enable cluster traces by running weka debug traces start.

DEBUG

TracesFreezePeriodActive

Trace freeze is active

After resolving the underlying issue, contact the Customer Success Team to deactivate the trace freeze if it persists.

DEBUG

UdpModePerformanceWarning

Backend container configured in UDP mode

If this is a misconfiguration, add network devices to the specified backend container using weka cluster container net add.

DEBUG

UnstableHosts

Host marked unstable during upgrade

Review container status and logs to identify and resolve hardware, network, or service problems before retrying the upgrade.

DEBUG

UnwritableDisksConfigured

A drive is set to unwritable

If the drive remains unwritable after maintenance, contact the Customer Success Team.

DEBUG

UserQuotasHardLimitReached

User quota hard limit exceeded

Run weka fs quota list --user to list users that have exceeded their hard quota limits. Free up space for these users or increase their hard quota limit.

WARNING

UserQuotasSoftLimitReached

User quota soft limit exceeded

Run weka fs quota list --user to list users that have exceeded their soft quota limits. Free up space for these users or increase their hard quota limit.

WARNING

WTracerDaemonWriteIOFailures

WTracer Daemon write I/O failures detected

Review System Events for details and check for storage or network issues affecting trace data writes.

WARNING

WTracerLostTraces

WTracer lost traces detected

Review System Events and check for high system load, memory pressure, or I/O bottlenecks.

MINOR

Last updated