Thursday, October 8, 2026

Instance Terminated When ORA-00494: enqueue [CF] held for too long by CKPT and Killed

Oracle Instance is being terminated due to CKPT holding enqueue CF too long and CKPT is killed by Oracle Background Process.
This Blog demonstrates one step-by-step test case.

Note: Tested in Oracle 19c with:


1. Test Steps


Open one UNIX window, and two Sqlplus sessions (SID-1 and SID-2).


1.1 UNIX Window


On UNIX Window, find ckpt process and its associated control files,
set GDB breakpoint when pread64 on them.
 
$ > ps -efl |grep ckpt
0 S oracle     81749       1  0  80   0 - 1246636 -    11:34 ?        00:00:00 ora_ckpt_testdb01

$ > lsof -p 81749 | grep control

ora_ckpt_ 81749 oracle  256uW  REG    8,1  32620544  671095912 /oratestdb01/oradata/testdb01/control01.ctl
ora_ckpt_ 81749 oracle  257uW  REG    8,1  32620544  671095913 /oratestdb01/oradata/testdb01/control02.ctl
ora_ckpt_ 81749 oracle  258uW  REG    8,1  32620544  671095914 /oratestdb01/oradata/testdb01/control03.ctl
Here gdb commands and output:
 
$ > gdb -p 81749

(gdb) display $rdi
  1: $rdi = 31
(gdb) break pread64 if $rdi == 256 || $rdi == 257 || $rdi == 258
  Breakpoint 1 at 0x7fc984de9870 (2 locations)
(gdb) info b
  Num     Type           Disp Enb Address            What
  1       breakpoint     keep y   
          stop only if $rdi == 256 || $rdi == 257 || $rdi == 258
  1.1                         y     0x00007fc984de9870 
  1.2                         y     0x00007fc9855052b0 
(gdb) c
  Continuing.

Breakpoint 1, 0x00007fc9855052b0 in pread64 () from /lib64/libpthread.so.0
  1: $rdi = 256
(gdb) info r
  rax            0x1                 1
  rbx            0x0                 0
  rcx            0x4000              16384
  rdx            0x4000              16384
  rsi            0x7fc9827a1000      140503454191616
  rdi            0x100               256


1. 2. Sqlplus SID-1


On the first Sqlplus session, run some DML and then force Oracle Database to perform a checkpoint.
CKPT starts to read control files, and stops on pread64 breakpoint.
Oracle v$session shows CKPT event "control file sequential read".
 
create table ksun_tab as 
select level x, rpad('ABC', 500, 'X') y from dual connect by level <= 1e4; 

Sqlplus > update ksun_tab set y = rpad('ABC', 500, 'x');

  10000 rows updated.

Sqlplus > commit;

  Commit complete.

Sqlplus > alter system checkpoint;
  ERROR:
  ORA-03114: not connected to ORACLE
  
  alter system checkpoint
  *
  ERROR at line 1:
  ORA-03113: end-of-file on communication channel
  Process ID: 80361
  Session ID: 283 Serial number: 3022


1.3. Sqlplus SID-2


On the second Sqlplus session, force Oracle Database to begin writing to a new redo log file group (log switch),
and Oracle Database also begins to perform a checkpoint, hence LGWR is blocked with Event: "enq: CF - contention":
 
Sqlplus > alter system switch logfile;
  ERROR:
  ORA-03114: not connected to ORACLE
  
  alter system switch logfile
  *
  ERROR at line 1:
  ORA-03113: end-of-file on communication channel
  Process ID: 81230
  Session ID: 24 Serial number: 25591


1.4. Instance is being terminated


After 900 seconds, alert.log shows:
 
Cause - 'Instance is being terminated due to ospid CKPT holding an enqueue CF'
Hiddne parameter "_controlfile_enqueue_timeout" controls file enqueue timeout in seconds with default 900 seconds.


2. DB alert.log


alert.log shows that CKPT is killed by one Background processes (e.g, LGWR, ARCn, TTnn, MMON) due to
ORA-00494: enqueue [CF] held for too long (more than 900 seconds) by CKPT.
 
Errors in file /orabin/app/oracle/admin/testdb01/diag/rdbms/testdb01/testdb01/trace/testdb01_mmon_80100.trc  (incident=33857):
ORA-00494: enqueue [CF] held for too long (more than 900 seconds) by 'inst 1, osid 80070'
Incident details in: /orabin/app/oracle/admin/testdb01/diag/rdbms/testdb01/testdb01/incident/incdir_33857/testdb01_mmon_80100_i33857.trc
2026-08-25T10:37:19.316424+02:00
Killing enqueue blocker (pid=80070) on resource CF-00000000-00000000-00000000-00000000 by (pid=80100)
 by killing session 273.24900
KILL SESSION for sid=(273, 24900):
  Reason = RAC enqueue blocker
  Mode = KILL SOFT -/-/-
  Requestor = MMON (orapid = 32, ospid = 80100, inst = 1)
  Owner = Process: CKPT (orapid = 21, ospid = 80070)
  Result = ORA-29
Killing enqueue blocker (pid=80070) on resource CF-00000000-00000000-00000000-00000000 by (pid=80100)
 by terminating the process
MMON (ospid: 80100): terminating the instance due to ORA error 2103
Cause - 'Instance is being terminated due to ospid 80070 holding an enqueue CF-00000000-00000000-00000000-00000000)'
2026-08-25T10:37:19.372998+02:00
System state dump requested by (instance=1, osid=80100 (MMON)), summary=[abnormal instance termination].


3 Background Incident and Trc files

 
Dump continued from file: /orabin/app/oracle/admin/testdb01/diag/rdbms/testdb01/testdb01/trace/testdb01_mmon_80100.trc
[TOC00001]
ORA-00494: enqueue [CF] held for too long (more than 900 seconds) by 'inst 1, osid 80070'

[TOC00001-END]
[TOC00002]
========= Dump for incident 33857 (ORA 494) ========
[TOC00003]
----- Beginning of Customized Incident Dump(s) -----
-------------------------------------------------------------------------------
ENQUEUE [CF] HELD FOR TOO LONG
 
enqueue holder: 'inst 1, osid 80070'
 
Process 'inst 1, osid 80070' is holding an enqueue for maximum allowed time.
The process will be terminated after collecting some diagnostics.
It is possible the holder may release the enqueue before the termination 
is attempted. If so, the process will not be terminated and 
this incident will only be served as a warning.
Oracle Support Services triaging information: to find the root-cause, look
at the call stack of process 'inst 1, osid 80070' located below. Ask the
developer that owns the first NON-service layer in the stack to investigate.
Common service layers are enqueues (ksq), latches (ksl), library cache
pins and locks (kgl), and row cache locks (kqr).

  [880 samples,                                            10:35:48 - 10:50:48]
    waited for 'control file sequential read', seq_num: 16046
      p1: 'file#'=0x0
      p2: 'block#'=0x1
      p3: 'blocks'=0x1
      
Problem Key: ORA 494
Error: ORA-494 [[CF]] [900] [1] [80070] [] [] [] [] [] [] [] []
[00]: dbgexExplicitEndInc [diag_dde]
[01]: dbgeEndDDEInvocationImpl [diag_dde]
[02]: dbgeEndSpltInvokOnRec [diag_dde]
[03]: dbgePostErrorKGE [diag_dde]
[04]: dbkePostKGE_kgsf [rdbms_dde]
[05]: kgeade []
[06]: kgeselv []
[07]: ksesec3 [KSE]
[08]: ksdx_cmdreq_wait_for_pending [VOS]<-- Signaling
[09]: ksdxdocmdmultex [VOS]


*Note: 10:35:48 - 10:50:48 is 15 minutes (900 seconds)
4. bpftrace CKPT io_submit
 
bpftrace -v -e 'tracepoint:syscalls:sys_enter_io_submit /strncmp("ora_ckpt", comm, 7) == 0/ {
              $iocb_sizeof = sizeof(struct iocb);
              $reqs = args->nr;
              $iocbpp_ptr = (struct iocb **)args->iocbpp;
              time("%H:%M:%S --- ");
              printf("%s, iocb_sizeof: %d, io_submit I/O request: %d\n", comm, $iocb_sizeof, $reqs);
              $i = 1; while ($i <= $reqs) { $iocbpp_ptr_2 = (struct iocb **)($iocbpp_ptr + $iocb_sizeof*($i-1));
                                            $iocbpp_struct = (struct iocb *)(*$iocbpp_ptr_2);
                                            printf("--- File: %d --- ptr_2 %p, aio_fildes: %d, aio_nbytes: %d, aio_offset: %ld, aio_lio_opcode(1=PWRITE): %d, aio_data:0X%X\n",
                                                   $i, $iocbpp_ptr_2, $iocbpp_struct->aio_fildes, $iocbpp_struct->aio_nbytes, $iocbpp_struct->aio_offset, $iocbpp_struct->aio_lio_opcode, $iocbpp_struct->aio_data);
                                            @All_CNT = count(); @CNT[$iocbpp_struct->aio_fildes] = count(); @SUM[$iocbpp_struct->aio_fildes] = sum( $iocbpp_struct->aio_nbytes);
                                            if ($i > 50) { break; }
                                            $i++ }
              printf("\n");}'