Tuesday, February 3, 2015

undefined reference to `nnfyboot'


When installing or relinking Oracle executables, it may be possible to see such error. The solution is :
1. Install addtitional packages (if not installed, 32bit packages such as glibc-32bit,glibc-devel-32bit,gcc-32bit and other required).
2. look through install.log and search messages kind of 'No such file or directory'. Try to fix errors (see item 1) and run all failed commands by hand in console, using oracle account.
For example, in following text

INFO: (if [ "compile" = "compile" ] ; then \
          echo "Building 32bit version of nnfgt.o"; \
          /opt/oracle/product/10gR2/db/bin/gennfgt > nnfgt.c ;\
          gcc -m32  -c nnfgt.c ;\
          rm -f /opt/oracle/product/10gR2/db/lib32/nnfgt.o ;\
          mv nnfgt.o /opt/oracle/product/10gR2/db/lib32/ ;\
          /usr/bin/ar rv /opt/oracle/product/10gR2/db/lib32/libn10.a /opt/oracle/product/10gR2/db/lib32/nnfgt.o ;\
          echo "Building 64bit version of nnfgt.o"; \
          /opt/oracle/product/10gR2/db/bin/gennfgt > nnfgt.c ;\
          gcc  -c nnfg
INFO: t.c ;\
          rm -f /opt/oracle/product/10gR2/db/lib/nnfgt.o ;\
          mv nnfgt.o /opt/oracle/product/10gR2/db/lib/ ;\
          /usr/bin/ar rv /opt/oracle/product/10gR2/db/lib/libn10.a /opt/oracle/product/10gR2/db/lib/nnfgt.o ; fi)

INFO: Building 32bit version of nnfgt.o

INFO: In file included from /usr/include/features.h:371,
                 from /usr/include/sys/types.h:27,
                 from nnfgt.c:7:
/usr/include/gnu/stubs.h:7:27: error: gnu/stubs-32.h: No such file or directory

INFO: mv: cannot stat `nnfgt.o': No such file or directory

INFO: /usr/bin/ar: /opt/oracle/product/10gR2/db/lib32/nnfgt.o: No such file or directory

INFO: Building 64bit version of nnfgt.o

INFO: r - /opt/oracle/product/10gR2/db/lib/nnfgt.o

The file gnu/stubs-32.h, needed to build 32-bit version of nnfgt.o, needed for construction of libclntsh.so, is absent. Install required package glibc-devel-32bit and run by hand above commands.

3. relink executables running relink command.

Friday, January 9, 2015

The active version of Oracle Clusterware is not 10g Release 2

Silent install of 10gR2 database on 12c or 11g clusterware environment fails with following error :

"The active version of Oracle Clusterware is not 10g Release 2"

Flags such as "-ignore..." don't help.

So the solution is to edit file /stage/prereq/db/db_prereq.xml and remove whole PREREQUISITE NAME="Detect10.2CRS" section, or remove its content,related to rules.

After that you'll be able to complete the installation. 

And don't forget to pin cluster nodes with:

/bin/crsctl pin css -n

Thursday, January 8, 2015

Hanged or stalled Grid Infrastructure (11g or 12c) installation performing remote operations on Linux

The installation process goes fine, but suddenly is's stalled in the middle (50-60%) for unknown reason. That's becouse java installation process unable to transmit copied data (files and dirs) to oracle processes named ractrans, running on remote nodes in listening mode and basically performing all work for filling up remote oracle homes.

In my case that was becouse of sysctl.conf settings.

The content I had :

net.core.rmem_default = 25165824
net.core.wmem_default = 25165824
net.core.rmem_max = 25165824
net.core.wmem_max = 25165824
net.core.rmem_max = 134217728
net.core.wmem_max = 134217728
net.ipv4.tcp_rmem = 134217728 134217728 134217728
net.ipv4.tcp_wmem = 134217728 134217728 134217728
net.core.netdev_max_backlog = 300000

The content I had to modify :

net.core.rmem_default = 4194304
net.core.wmem_default = 262144
net.core.rmem_max = 4194304
net.core.wmem_max = 262144
net.ipv4.tcp_rmem = 4096        87380   174760
net.ipv4.tcp_wmem = 4096        16384   131072
net.core.netdev_max_backlog = 1000

Honestly, I suspect, only net.ipv4.tcp_rmem and net.ipv4.tcp_wmem parameters play main role in this issue. Just in case I decided to set default values for other in the list above.

I.e., decreasing amount of memory per tcp socket and buffer space, slowing ethernet speed, took results.

P.S.

1. Just as surveillance. Oracle Universal Installer creates six (6) process per remote node to sync contents of grid_home. If you perform installation only for two nodes, you could not face with this issue.

2. It generally depends on speed of your public network stack. If you have fast enough switches and adapters, probably you will not bump into it.

3. It also depends on speed your storage (including mounting options). Any measure decreasing speed of copying files, can help.

Saturday, November 8, 2014

How to modify interconnect device/settings in 11gR2 Clusterware

Let's imagine the following situation.
You have two interconnect interfaces. Suddenly, all of them have become unaccessible, your rac immediately stopped. How to rule it out ?
Build up new interconnect, using new cards, of course.
But you will not be able to run clusterware services like HAIP and other becouse your new interconnect device name does not consistent with interconnect device names in OLR (Oracle Local Registry on each node) and OCR.
Actions ?

1. Rename you interconnect device by operating system means (modify udev net-persistent-device-names file, for example, and run udevadm trigger command on each node), and rerun your clusterware afterwards.

OR

2. Modify OLR content on each node

$ cd $CRS_HOME/gpnp/$(hostname)/profiles/peer

Here will be signed profile.xml, modifications only allowed with gpnptool ($CRS_HOME/bin) utility.

query profile.2 in some manner like that:
$ gpnptool getpval -p=profile.2 -net
$ gpnptool getpval -p=profile.2 -net2:net_ip
$ gpnptool getpval -p=profile.2 -net2:net_ada

unsign profile.xml to new profile.2 :
$ gpnptool unsign -p=profile.xml -o=profile.2

rename interconnect adapter name in profile.2 :

$ gpnptool edit -p=profile.2 -net2:net_ada=new_if_name -o=profile.2 -ovr

set if needed new IP subnet for interconnect :
$ gpnptool edit -p=profile.2 -net2:net_ip=subnet.0 -o=profile.2 -ovr

sign profile.2 :
$ gpnptool sign -p=profile.2 -o=profile.2.xml.signed -w=/$CRS_HOME/gpnp/$(hostname)/wallets/peer

make a rotation :
$ mv profile.xml profile.xml.orig

$ cp profile.2.xml.signed profile.xml

run clusterware :
# /etc/init.d/ohasd start

After successful running, with help of oifcfg, delete old interconnect data from OCR by
$ oifcfg delif -global old_if_name.

Repeat above steps related to profile.xml on each node

Tuesday, September 23, 2014

Updating FabricOS using scp - The server is inaccessible or firmware path is invalid.

If you're getting something like this during updating brocade san switch firmware:
 
Failed to access scp://oracle:**********@hostname//vol200/temp/hp_hba/brocade/fos_721b/v7.2.1b/release.plist
The server is inaccessible or firmware path is invalid. Please make sure the server name/IP address and the firmware path are valid, the protocol and authentication are supported. It is also possible that the RSA host key could have been changed and please contact the System Administrator for adding the correct host key.

than you have two options :

1. Use ftp server to download firmware (I tried with vsftpd, but having trouble with binary/ascii mode having completely disabling ascii in the config I've left it).

2. Modify sshd_config, because firmwaredownload over scp/sftp require plain ssh password authentication

PasswordAuthentication yes

Reboot ssh server and have a lot of fun.
P.S.  Avoid ssh passwords with spaces for firmwaredownload.

Monday, September 8, 2014

How to free up space, "used" by deleted file which is held by running application

1. Just kill your application

2. If you can't kill the application, you can :

a) figure out the pid of the app;
b) figure out deteled filename (lsof | grep deleted | grep pid)
c) run

# : > /proc/$app_pid/fd/$decrpiptor_corresponding_to_deleted_file


Wednesday, August 13, 2014

How to unmount stale nfs share

If you suddenly have got stale and unable to read nfs mount point becouse of changed IP of the nfs server or something else, you should use following to unmount nfs share and after that remount it.

1.

# umount -lf /nfs_share

2. or

# umount.nfs /nfs_share -l -f

It should work