How to Diagnose and Troubleshoot Node Evictions in Oracle RAC Environment

How to Diagnose and Troubleshoot Node Evictions in Oracle RAC Environment

Oracle RAC is built to ensure that multiple database instances work together seamlessly for high availability. If Oracle Clusterware finds that a node is unhealthy or has trouble communicating with the rest of the cluster, it may remove the node.

Removing a node is a significant event in RAC because Oracle Clusterware does this to protect the cluster from issues like split-brain situations, network failures, problems with voting disks, or lack of resources.

What Is Oracle RAC Node Eviction?

In simple terms, node eviction is when Oracle Clusterware takes a RAC node out of the cluster because it thinks the node is unhealthy or cannot communicate properly with the cluster.

For example, imagine a three-node RAC cluster:

  • Node 1 – ONLINE
  • Node 2 – ONLINE
  • Node 3 – ONLINE

If Node 2 stops responding to Cluster Synchronization Service (CSS) heartbeats, Oracle Clusterware may determine that Node 2 is unhealthy and evict it from the cluster.

After the problem is resolved, the evicted node can normally rejoin the cluster.

Why Does Oracle Evict a RAC Node?

The primary purpose of eviction is cluster protection.

Oracle RAC must prevent a situation where two parts of the cluster believe they are independently active. This is commonly associated with the concept of split-brain.

Oracle Clusterware constantly checks cluster communication and key resources. If a node can’t communicate reliably in the set timeout, Clusterware may act to protect the other nodes in the cluster.

Common Causes of RAC Node Eviction

  • Interconnect/network problems
  • Packet loss between RAC nodes
  • Network interface failures
  • Network congestion or latency
  • Voting disk I/O problems
  • Slow or failed voting disk access
  • Storage latency
  • CPU starvation
  • Very high CPU utilization
  • Run-queue pressure
  • Operating system resource starvation
  • Memory pressure
  • Heavy swapping
  • Operating system hangs
  • Oracle Grid Infrastructure or Clusterware problems

Important RAC Parameters for Node Eviction

Two important CSS parameters that every Oracle RAC DBA should understand are:

  • misscount – related to the time allowed for missed CSS network heartbeats.
  • disktimeout – related to the time allowed for voting disk I/O.

Oracle’s CRSCTL documentation states that the default values are around 30 seconds for misscount and 200 seconds for disktimeout, but you should always verify the actual settings in the RAC environment being troubleshot.

6 Key Checks to Diagnose RAC Node Eviction

1. Check CSS Misscount

The first check is the CSS misscount value.

crsctl get css misscount

This command retrieves the configured CSS misscount value. Oracle documents the syntax as: crsctl get css parameter.

A typical output may look similar to:

CRS-4678: Successful get misscount 30 for Cluster Synchronization Services.

The important point is that you should not immediately increase misscount just because a node was evicted. First determine why CSS heartbeats were missed.

2. Check CSS Disk Timeout

The second important check is the voting disk timeout.

crsctl get css disktimeout

A typical result may be:

CRS-4678: Successful get disktimeout 200 for Cluster Synchronization Services.

Oracle documents disktimeout as a CSS parameter associated with voting disk communication. The documented default is 200 seconds.

If the eviction is related to voting disk or storage problems, investigate the storage layer rather than simply increasing the timeout.

3. Check Clusterware and Operating System Logs

Logs are one of the most important sources of information when investigating a RAC node eviction.

Start by checking the Clusterware logs and operating system logs, for example:

/var/log/messages

Depending on the Linux distribution, the relevant operating system logs may be located in different files or collected through systemd/journald.

Also investigate the Oracle Grid Infrastructure/Clusterware logs around the exact time of the eviction.

Look for messages related to:

  • CSS
  • CRS
  • Voting disk
  • Network heartbeat
  • Missed heartbeat
  • Node failure
  • Reconfiguration
  • Network interface errors
  • I/O timeout
  • Resource starvation

4. Use OCLUMON to Investigate CPU and Memory Pressure

oclumon is extremely useful when investigating whether operating-system resource pressure contributed to a RAC node eviction.

For example:

oclumon dumpnodeview -cpu

You can also examine process information:

oclumon dumpnodeview -process

Disk information can be investigated with:

oclumon dumpnodeview -device

How to Diagnose and Troubleshoot Node Evictions in Oracle RAC Environment

Oracle RAC is designed to provide high availability by allowing multiple database instances to work together as a cluster. However, when Oracle Clusterware detects that a node is unhealthy or cannot reliably communicate with the rest of the cluster, it may evict the node.

Node eviction is a serious RAC event because Oracle Clusterware performs it to protect the cluster from situations such as split-brain, network failures, voting disk problems, or severe resource starvation.

What Is Oracle RAC Node Eviction?

In simple terms, node eviction means Oracle Clusterware removes a RAC node from the cluster because it believes that the node is no longer healthy or cannot maintain reliable cluster communication.

For example, imagine a three-node RAC cluster:

  • Node 1 – ONLINE
  • Node 2 – ONLINE
  • Node 3 – ONLINE

If Node 2 stops responding to Cluster Synchronization Service (CSS) heartbeats, Oracle Clusterware may determine that Node 2 is unhealthy and evict it from the cluster.

After the problem is resolved, the evicted node can normally rejoin the cluster.

Why Does Oracle Evict a RAC Node?

The primary purpose of eviction is cluster protection.

Oracle RAC must prevent a situation where two parts of the cluster believe they are independently active. This is commonly associated with the concept of split-brain.

Oracle Clusterware continuously monitors cluster communication and important resources. If a node cannot communicate reliably within the configured timeout period, Clusterware may take action to protect the remaining cluster.

Common Causes of RAC Node Eviction

  • Interconnect/network problems
  • Packet loss between RAC nodes
  • Network interface failures
  • Network congestion or latency
  • Voting disk I/O problems
  • Slow or failed voting disk access
  • Storage latency
  • CPU starvation
  • Very high CPU utilization
  • Run-queue pressure
  • Operating system resource starvation
  • Memory pressure
  • Heavy swapping
  • Operating system hangs
  • Oracle Grid Infrastructure or Clusterware problems

Important RAC Parameters for Node Eviction

Two important CSS parameters that every Oracle RAC DBA should understand are:

  • misscount – related to the time allowed for missed CSS network heartbeats.
  • disktimeout – related to the time allowed for voting disk I/O.

Oracle’s current CRSCTL documentation lists the default values as approximately 30 seconds for misscount and 200 seconds for disktimeout, although the actual configuration should always be checked on the RAC environment being troubleshot.

6 Key Checks to Diagnose RAC Node Eviction

1. Check CSS Misscount

The first check is the CSS misscount value.

crsctl get css misscount

This command retrieves the configured CSS misscount value. Oracle documents the syntax as: crsctl get css parameter.

A typical output may look similar to:

CRS-4678: Successful get misscount 30 for Cluster Synchronization Services.

The important point is that you should not immediately increase misscount just because a node was evicted. First determine why CSS heartbeats were missed.

2. Check CSS Disk Timeout

The second important check is the voting disk timeout.

crsctl get css disktimeout

A typical result may be:

CRS-4678: Successful get disktimeout 200 for Cluster Synchronization Services.

disktimeout as a CSS parameter associated with voting disk communication. The documented default is 200 seconds.

If the eviction is related to voting disk or storage problems, investigate the storage layer rather than simply increasing the timeout.

3. Check Clusterware and Operating System Logs

Logs are one of the most important sources of information when investigating a RAC node eviction.

Start by checking the Clusterware logs and operating system logs, for example:

/var/log/messages

Depending on the Linux distribution, the relevant operating system logs may be located in different files or collected through systemd/journald.

Also investigate the Oracle Grid Infrastructure/Clusterware logs around the exact time of the eviction.

Look for messages related to:

  • CSS
  • CRS
  • Voting disk
  • Network heartbeat
  • Missed heartbeat
  • Node failure
  • Reconfiguration
  • Network interface errors
  • I/O timeout
  • Resource starvation

4. Use OCLUMON to Investigate CPU and Memory Pressure

oclumon is extremely useful when investigating whether operating-system resource pressure contributed to a RAC node eviction.

For example:

oclumon dumpnodeview -cpu

You can also examine process information:

oclumon dumpnodeview -process

Disk information can be investigated with:

oclumon dumpnodeview -device

Network information can be investigated with:

oclumon dumpnodeview -nic

For a more complete node view, you can use:

oclumon dumpnodeview -v

oclumon dumpnodeview as a way to view Cluster Health Monitor information in node-view form. It can expose system, CPU, process, device, network interface, filesystem and other metrics depending on the options used

5. Investigate CPU, Memory and Swap Activity

A node can appear to have a network or CSS problem when the underlying issue is actually severe operating-system resource pressure.

Check:

  • CPU utilization
  • CPU run queue
  • System load
  • Free memory
  • Swap utilization
  • Disk I/O wait
  • Network errors
  • Network retransmissions

For example, if the operating system is heavily overloaded, important Oracle Grid Infrastructure processes may not receive CPU time quickly enough. This can contribute to missed heartbeats and ultimately node eviction.

6. Check the AHF/CHM Diagnostic Information

If Autonomous Health Framework (AHF) is available in the environment, use its diagnostic capabilities to collect and analyze information around the eviction.

The Cluster Health Monitor information can be particularly useful because it provides historical operating-system metrics that help correlate an eviction with CPU, memory, network, and I/O activity.

For example:

oclumon dumpnodeview -last "01:00:00"

This can be used to inspect recent node metrics. Oracle’s current AHF documentation provides additional options for filtering, sorting and querying historical node metrics.

How RAC Node Eviction Works – Simplified

The process can be understood using this simplified flow:

  1. RAC nodes continuously communicate with each other.
  2. CSS monitors cluster communication.
  3. A node starts missing required heartbeats.
  4. Clusterware determines that the node may be unhealthy.
  5. Clusterware performs the required cluster reconfiguration.
  6. The unhealthy node is evicted to protect the cluster.
  7. The node can restart/rejoin after the underlying problem is resolved.

Example: Network Heartbeat Failure

Suppose you have a three-node RAC cluster:

NODE1
NODE2
NODE3

Now imagine that the private interconnect between NODE2 and the other RAC nodes starts dropping packets.

The sequence could look like this:

Network packet loss
↓
CSS heartbeats are missed
↓
Clusterware detects communication failure
↓
Node becomes unhealthy
↓
Clusterware performs eviction/reconfiguration
↓
NODE2 leaves the cluster
↓
Problem is fixed
↓
NODE2 rejoins the RAC cluster

Example: Voting Disk I/O Problem

Another common scenario is a storage problem.

Suppose the voting disk is hosted on shared storage and the storage system suddenly experiences very high latency.

The sequence could be:

Storage latency increases
↓
Voting disk I/O becomes slow
↓
CSS cannot complete required voting disk operations
↓
Clusterware detects the problem
↓
Node may be evicted
↓
RAC cluster reconfiguration occurs

In this situation, the DBA should investigate the storage path, SAN, ASM disk latency, multipathing and underlying infrastructure rather than simply changing CSS timeout values.

Important Commands Cheat Sheet

PurposeCommand
Check CSS misscountcrsctl get css misscount
Check voting disk timeoutcrsctl get css disktimeout
Check CPU metricsoclumon dumpnodeview -cpu
Check process metricsoclumon dumpnodeview -process
Check disk metricsoclumon dumpnodeview -device
Check network metricsoclumon dumpnodeview -nic
Verbose node viewoclumon dumpnodeview -v
Recent node metricsoclumon dumpnodeview -last "01:00:00"
Check OS messages/var/log/messages

What NOT to Do After a Node Eviction

One of the biggest mistakes during RAC troubleshooting is changing timeout parameters without first finding the root cause.

Do not immediately increase misscount just to stop future evictions.

If the real problem is:

  • Network packet loss
  • Bad NIC
  • Network congestion
  • High CPU utilization
  • Memory pressure
  • Swap activity
  • Slow storage
  • Voting disk latency
  • Operating system problems

then increasing the timeout may only hide the real problem and potentially make detection of an actual failure slower.

Recommended RAC Node Eviction Troubleshooting Approach

When a production RAC node is evicted, follow a structured approach.

  1. Record the exact eviction timestamp.
  2. Identify the affected node.
  3. Check Clusterware/CSS logs around that timestamp.
  4. Check the operating system logs.
  5. Check the private interconnect and network interfaces.
  6. Check voting disk and ASM/storage latency.
  7. Check CPU and memory utilization.
  8. Check swap and I/O wait.
  9. Use OCLUMON/AHF historical metrics to correlate the event.
  10. Determine whether the root cause is network, storage, OS or Grid Infrastructure.
  11. Fix the underlying issue.
  12. Monitor the cluster for additional evictions.

This entry was posted in Oracle on by .
Unknown's avatar

About SandeepSingh

Hi, I am working in IT industry with having more than 15 year of experience, worked as an Oracle DBA with a Company and handling different databases like Oracle, SQL Server , DB2 etc Worked as a Development and Database Administrator.

Leave a Reply