Repository navigation
VR Loses Instance IP Address #3550
Description
Activity
Can you review/test and advise @nvazquez and if it's an easy fix we can aim to get it in 4.13.0.0
From which versions did you upgrade @div8cn ?
Hi @div8cn I could not reproduce this on 4.11.3. Did you hit the same issue on prior existing networks as well as on new networks after upgrading or just the existing ones?
I am having problems after upgrading VR after upgrading 4.11.3.
I can't reproduce this problem in my test environment right now.
There are currently 8 share networks in my production environment, all of which have different levels of IP loss, and both occur after upgrading VR.
From which versions did you upgrade @div8cn ?
I upgraded from 4.11.2 to 4.11.3
The host is xenserver 7.1.2
Thanks @div8cn, did you hit the issue recreating the VRs using the new system VM template?
System template has been updated
Vr has also completed the upgrade
Couldn't reproduce myself...
Couldn't reproduce either
Thanks @andrijapanicsb @nvazquez I've moved it to next minor milestone, this is not a blocker for 4.13.
@andrijapanicsb @nvazquez @rhtyd We have similar bug after upgrade from 4.11.2 to 4.11.3 (kvm hosts)
We are managing a 800+ VM environment with XenServer 7.1CU2 and ACS 4.11.3. We have hit this instance IP lose issue for multiple times which has not been seen while we stayed with ACS 4.11.2.
@DmB991 Just curious, how many running VMs in your environment which you saw this issue ? And did you take any chance to rebuild the VR if it solved this issue ?
@luhaijiao We use CloudStack for testing our network product, so we create and destroy huge amount of networks and VMs(more than 5000) every day and we see this issue really often. We just returned to version 4.11.2
Can you guys perhaps make a correlation with the out-of-band live migration (vSphere, DRS, XenServer,etc. - i.e. that ACS is not aware of) - and if this triggers the thing.
We are observing specific situation which seems related to DRS doing the out of band migration, then that VM's power off report is sent back to ACS, so it immediately removes it's DNS/DHCP (new in 4.11.3, to clean the garbage - it's bit complicated to explain atm) - but later ACS gets another power on report - so doesn't issue specifically an internal "Stop" command - and VM remains happily running, but it's DNS/DHCP records are gone...
We are planning some changes in the logic, i.e. wipe DNS/DHCP only when the VM is expunged, not when it's stopped etc.
But I would be happy if you can make a correlation of your to mine situation/root cause.
Thank you for your reply
Sometimes we do skip ACS and migrate VMs directly through xencenter.
We will test again, hoping to reproduce the fault.Fixed in master @div8cn @DmB991 @luhaijiao - we now remove DHCP config only when VM is expunged - not before - so any DRS/live migrations outside of the CloudStack - will not influence VM to lose its IP address. Please test when you have time.
@andrijapanicsb Can you link the relevant PR? I couldn't find the code fixing this in master and also in 4.11. Thanks in advance
This looks like is closed, please confirm and close @div8cn @andrijapanicsb @nvazquez
Closed and fixed in master and 4.13.1.
ISSUE TYPE
COMPONENT NAME
CLOUDSTACK VERSION
CONFIGURATION
OS / ENVIRONMENT
SUMMARY
After the ACS is upgraded to 4.11.3, the VM on the shared network usually loses its IP address (it cannot obtain an IP address in the guest VM).
Analysis of the cloud.log in vr, we found that sometimes restarting or deleting the A VM will cause the B VM to disappear in /etc/dhcphosts.txt.
Because /etc/dhcphosts.txt is inconsistent with the data in /var/lib/misc/dnsmasq.leases. This triggers delete_leases to remove the IP address of the B VM from the DHCP server.
This bug should come from #3351
STEPS TO REPRODUCE
EXPECTED RESULTS
ACTUAL RESULTS