Mabuhay

Hello world! This is it. I've always wanted to blog. I don't want no fame but just to let myself heard. No! Just to express myself. So, I don't really care if someone believes in what I'm going to write here nor if ever someone gets interested reading it. My blogs may be a novel-like, a one-liner, it doesn't matter. Still, I'm willing to listen to your views, as long as it justifies mine... Well, enjoy your stay, and I hope you'll learn something new because I just did and sharing it with you.. Welcome!
Showing posts with label reboot. Show all posts
Showing posts with label reboot. Show all posts

Sunday, June 29, 2008

Reboot after panic: Data page fault

One of the servers that we monitor rebooted on a panic. The dumps, I think, will be submitted to the HPRC for decoding what caused the panic. But, for DPF's, it is most unlikely that it was caused by hardware failure. Probably, some application passed something on the kernel that it didn't know how to process it. Anyway, just an opinion. I'm not sure if I can get anything from the DTS. A root cause analysis is needed for this.

Here are the files on the crash dumps.

[root@box1:/var/adm/crash/crash.0]
# ll
total 2113190
-rw-r--r-- 1 root root 1556 Jun 28 19:05 INDEX
-rw-r--r-- 1 root root 281784 Jun 28 19:02 SEOS
-rw-r--r-- 1 root root 134189056 Jun 28 19:02 image.1.1
-rw-r--r-- 1 root root 134197248 Jun 28 19:02 image.1.2
-rw-r--r-- 1 root root 134152192 Jun 28 19:03 image.1.3
-rw-r--r-- 1 root root 89419776 Jun 28 19:03 image.1.4
-rw-r--r-- 1 root root 134180864 Jun 28 19:04 image.2.1
-rw-r--r-- 1 root root 134180864 Jun 28 19:04 image.2.2
-rw-r--r-- 1 root root 134193152 Jun 28 19:04 image.2.3
-rw-r--r-- 1 root root 134168576 Jun 28 19:05 image.2.4
-rw-r--r-- 1 root root 36978688 Jun 28 19:05 image.2.5
-rw-r--r-- 1 root root 16007272 Jun 28 19:02 vmunix

[root@box1:/var/adm/crash/crash.0]
# more INDEX
comment savecrash crash dump INDEX file
version 2
hostname box1
modelname 9000/800/N4000-44
panic Data page fault
dumptime 1214693427 Sat Jun 28 18:50:27 EDT 2008
savetime 1214694120 Sat Jun 28 19:02:00 EDT 2008
release @(#) $Revision: vmunix: vw: -proj selectors: CUPI80_BL2000_1108 -c 'Vw for CUPI80_BL2000_1108 buil
d' -- cupi80_bl2000_1108 'CUPI80_BL2000_1108' Wed Nov 8 19:24:56 PST 2000 $
memsize 4294967296
chunksize 134217728
module /stand/vmunix vmunix 16007272 3341476060
module /stand/dlkm/mod.d/SEOS SEOS 281784 3144992042
image image.1.1 0x0000000000000000 0x0000000007ff9000 0x0000000000000000 0x0000000000008987 356976270
image image.1.2 0x0000000000000000 0x0000000007ffb000 0x0000000000008988 0x0000000000015447 637338283
image image.1.3 0x0000000000000000 0x0000000007ff0000 0x0000000000015448 0x0000000000066d7f 35130470
image image.1.4 0x0000000000000000 0x0000000005547000 0x0000000000066d80 0x000000000007ffff 3529390633
image image.2.1 0x0000000000000000 0x0000000007ff7000 0x0000000000180000 0x00000000001a7307 719265648
image image.2.2 0x0000000000000000 0x0000000007ff7000 0x00000000001a7308 0x00000000001bec17 3529725656
image image.2.3 0x0000000000000000 0x0000000007ffa000 0x00000000001bec18 0x00000000001ca9df 560273249
image image.2.4 0x0000000000000000 0x0000000007ff4000 0x00000000001ca9e0 0x00000000001f5e0f 3332528375
image image.2.5 0x0000000000000000 0x0000000002344000 0x00000000001f5e10 0x00000000001fffff 3748535493

[root@box1:/var/adm/crash/crash.0]
# uptime
8:29pm up 1:28, 2 users, load average: 0.14, 0.18, 0.11

[root@box1:/var/adm/crash/crash.0]
# date
Sat Jun 28 20:29:31 EDT 2008

[root@box1:/var/adm/crash/crash.0]
# more /etc/shutdownlog
10:02 Mon Oct 10, 2005. Reboot:
11:22 Mon Oct 10, 2005. Reboot: (by SAM)
11:25 Mon Oct 10, 2005. Reboot: (by bdhp4420!root)
06:16 Tue Oct 11, 2005. Reboot: (by bdhp4420!root)
09:56 Thu Oct 20, 2005. Reboot: (by bdhp4420!root)
10:06 Thu Apr 20, 2006. Reboot: (by SAM)
10:12 Thu Apr 20, 2006. Reboot: (by bdhp4420!root)
19:05 Sat Jun 28 2008. Reboot after panic: Data page fault

[root@box1:/var/tombstones]
# ll -rt
total 252
-rw-r--r-- 1 root root 14720 Oct 7 2005 ts93
-rw-r--r-- 1 root root 14720 Oct 10 2005 ts94
-rw-r--r-- 1 root root 14720 Oct 10 2005 ts95
-rw-r--r-- 1 root root 14720 Oct 11 2005 ts96
-rw-r--r-- 1 root root 14720 Oct 20 2005 ts97
-rw-r--r-- 1 root root 14720 Apr 20 2006 ts98
-rw-r--r-- 1 root root 14720 Jun 28 19:02 ts99
-rw-r--r-- 1 root root 20873 Jun 28 19:08 cpumap

[root@box1:/var/tombstones]
#


On the side note, some application processes are not running. I saw earlier that one of them, Oracle, was started manually. So, I guess the same goes with the rest.

UPDATE: The team already opened an HPRC case for this and at the same time, a Problem Record (PR#54269).

Sunday, May 25, 2008

Reconfiguring an HP-UX (11iv1) kernel

It's whole new experience [heaven]. I thought it will remain a wish. Performing this is a long-shot in our working environment. But, what could be sweeter than a wish coming true? Nada! Of course, other than.. nah, nevermind! Ok, we have a project to have a kernel parameter changed. Upon doing a prework, we have this initial value:

[root@hpux09:/stand/build]
# kmtune -q vx_maxlink
Parameter Current Dyn Planned Module Version
==========================================
vx_maxlink 32767 - 32767

[root@hpux05:/stand/build]

We usually perform these kind of changes via SAM. But, wait! What the f*%$? Where's the vx_maxlink parameter?!? "Uhmm, hey [referring to my colleague], can you check it with L3?" So that's it. There are some parameters that does not show [or linked] on SAM. Changes to be made are to be performed via CLI. So, here are the steps I followed:

[root@hpux09:/stand/build]
# /usr/lbin/sysadm/system_prep -s system

[root@hpux09:/stand/build]
# kmtune -q vx_maxlink
Parameter Current Dyn Planned Module Version
==========================================
vx_maxlink 32767 - 32767

[root@hpux09:/stand/build]
# kmtune -s vx_maxlink=65534

[root@hpux09:/stand/build]
# kmtune -q vx_maxlink
Parameter Current Dyn Planned Module Version
==========================================
vx_maxlink 32767 - 65534

[root@hpux09:/stand/build]
# which mk_kernel
/usr/sbin/mk_kernel

[root@hpux09:/stand/build]
# mk_kernel -s system
Generating module: krm...
Generating module: SEOS...
Compiling conf.c...
Loading the kernel...
Generating kernel symbol table...

[root@hpux09:/stand/build]
# kmupdate

Kernel update request is scheduled.

Default kernel /stand/vmunix will be updated by
newly built kernel /stand/build/vmunix_test
at next system shutdown or startup time.


[root@hpux09:/stand/build]
# shutdown -ry 0
Shutdown cannot be run from a mounted file system -- exiting shutdown.
Change directories to the root volume ("/" will work) and try again.

[root@hpux09:/stand/build]
# cd /

[root@hpux09:/]
# shutdown -ry 0

SHUTDOWN PROGRAM
05/24/08 23:40:31 EDT

Broadcast Message from root (pts/5) Sat May 24 23:40:31...
PLEASE LOG OFF NOW ! ! !
System maintenance about to begin.
All processes will be terminated in 0 seconds.

Broadcast Message from root (pts/5) Sat May 24 23:40:31...
SYSTEM BEING BROUGHT DOWN NOW ! ! !

/sbin/auto_parms: DHCP access is disabled (see /etc/auto_parms.log)



For now, we have to wait for box to come up and check if the change we applied took effect. [Cross-finger] Hoping it did.?! It's driving me crazy [and very excited!].

What the f*%$ have I done?? I am doomed! The kernel parameter didn't change at all. The value is still under planned. My heart raced and pounded. Hey! I'm no superman. Looking for a reason to have the window time extended. Deym! [Temporary, still to hear a LOT about this during our weekly meeting] Fortunate for me (?), the change was so important that it left no choice for the application team, requestor, and box owner to extend the time and allow me to give it another go. But, this time? I got L3's attention! I consulted them, and they gave me an SOP [btw, for the record, of which I'm not aware of and was not provided]. [Another] But, to make sure, I let the L3 do it [I got my hands tied already, so I'm not taking any chances - not now but, definitely will love to do it again, anytime, anywhere!], and check how he did it a bit later. A few, very long, minutes later, he's working his magic. And here's how:

cd /stand/build
ll system [optional but essential]
kmtune -q vx_maxlink
/usr/lbin/sysadm/system_prep -s system
kmtune -s vx_maxlink=65534 -S ./system # This is what I missed; writing to system file
/usr/sbin/mk_kernel -s system
kmupdate
cd ..
cp -p system system_prev
mv build/system .
kmtune -q vx_maxlink
[now all I need is to reboot the box, and done!]

Well, folks, I hope you learned new. For me? I learned a TON!

And oh, make sure to watch out for the /stand FS getting full. You might end up just like it. I tell you, it's nasty. May be giving a system_prep will clear it... Well, just a thought. Good luck to us all.


This was added a bit later [July something of 2008].
Here's the procedure for HPUX 11.23: (explanation? Later)

# mv /stand/system /stand/system.orig
# kconfig -e /stand/system

CEdit /stand/system file and remove all "Tunable Parameters".
Copy and paste Tunable parameters ( lined between START KERNEL PARM and END KERNEL PARM) from /tmp/logfile.
Save and exit.

# kconfig -i /stand/system
# shutdown -r 0

Saturday, April 05, 2008

[Part of] 24 hour challenge over [halo-halo] senseless topics

I'm just starting this not knowing where it'll bring me nor even complete it. Not even sure if it will make sense at all.

I'm a bit off than the usual since I'm challenging myself to go 24H, online. Lucky for me, I'm with my happy colleagues. I like this, well, going to enjoy I guess. Hmmm, challenging night, I mean morning! Approximately 5 hours left before the sun rises. A new day starts as my day ends. Sleep seems to be rare lately. Is it worth it? Well, it depends. I respect those who really love their work BUT do you think you're productive? I mean, can you sustain your focus over a period of time? There may be exceptions but generally? I don't think so!

Going to I-don't-know-where, I'm thinking what else to write .......
.......
.......
.......
.......
.......
.......
.......
.......

It's 1:28 SGT now.

[Not so] Strange as it is, what happens when system reboot/crash/panic? I think I found my reason to stay awake for few hours.

++ Checking...
++ Still checking...

An HP-UX system crash is an unusual event. When a system panic occurs, it means that HP-UX encountered a condition that it did not know how to handle (or could not handle).

When the system crashes, HP-UX tries to save the image of physical memory, or certain portions of it, to predefined locations called dump devices. Then, when you next reboot the system, a special utility copies the memory image from the dump devices to the HP-UX file system area.

Prior to HP-UX Release 11.0, devices to be used as dump devices had to be defined in the kernel configuration, and they still can be. However, beginning with Release 11.0, a new, more flexible method for defining dump devices is available.

There are now multiple ways that dump devices can be configured. Here are three commonly used ways to define dump devices:

* In the kernel (as with releases prior to Release 11.0)
* During system initialization when the initialization script for crashconf runs (and reads entries from the /etc/fstab file)
* During run time, by an operator or administrator manually running the /sbin/crashconf command.

The dump process exists so that you have a way of capturing what your system was doing at the time of a crash. This is not for recovery purposes; processes cannot resume where they left off, following a system crash. Rather this is for analysis purposes, to help you determine why the system crashed in order to prevent it from happening again.

If you want to be able to capture the memory image of your system when a crash occurs (for later analysis), you need to define in advance the location(s) where HP-UX puts that image at the time of the crash. This location can be on local disk devices, or logical volumes.

Wherever you decide that HP-UX should put the dump, it is important to have enough space at that location (see “How Much Dump Space Do I Need?”) If you do not have enough space, not every page will be saved and you might not capture the part of memory that contains the instruction or data that caused the crash. If necessary, you can define more than one dump device so that if the first one fills up, the next one is used to continue the dumping process until the dump is complete or no more defined space is available. To guarantee that you have enough dump space, define a dump area that is at least as big as your computer’s physical memory, plus one megabyte. If you are doing a selective dump (which is the default dump mode in most cases), much less dump space will actually be needed. Full dumps require dump space equal to the size of your computer’s memory plus a little extra for header information.

In HP-UX Release 11i compressed dumps are enabled by default. However, dump compression will only occur if conditions in the crash environment are favorable. Do not plan your dump storage space based on potential compression but allow enough space for an uncompressed full or selective dump. See “Compressed Dump(HP-UX version 1 (B.11.11) or later)”.

Re: http://docs.hp.com/en/B2355-90950/ch05s05.html#bajcidic

++ I guess this will suffice. For now...

World Clock