Peripheral interrupt priority change due to project migration from SDK13.0 to SDK15.3

Hello,

after several problems resolved we were quite confident, that the migration to SDK15.3 was successfully completed, but when testing intensively we found a hopefully last problem which we don't know from SDK13.0. For certification purposes we need a stack (S140) with QDID, so we have migrated the project to SDK15.3...

I suppose that the problem is quite similar to the one described in case 223057, due to this I have set the define of APP_TIMER_CONFIG_IRQ_PRIORITY to 6, but unfortunately no change in behavior.

The problem appears after a certain while; I just explain what we do: At our nRF52840 we have connected an acceleration sensor using a SPI-Interface. This acceleration sensor creates an external HW signal when its internal buffer is full; this external interrupt is connected to a GPIO from the nRF52840. Additionally we have connected 2 external Hall signals at other GPIOs - the 1st for wake-up and speed measurement purposes, the second is only used for less than 1 second for measuring a time difference to the first one. In our test we create at this first Hall sensor for a period of 30 seconds a frequency of about 107 Hz which means every around 9 ms. At wake-up the acceleration sensor is configured to measuring mode and starts measurements; when buffer is full the nRF52840 reads that data block of 84 accelerations tripels for x-, y- and z-axis, calculates the length of the effective acceleration vector and stores the highest one. This one is then compared to the highest value of the next data block to get the overall highest one - perhaps this is quite a lot to do for an interrupt function, but in SDK13 it worked fine... The acceleration sensor starts automatically and creates the next buffer full interrupt; this is done until 1 second is achieved; then the procedure of measuring is restarted and the overall highest valie is reset to 0. We see that in one second up to 154 data blocks are read, calculated and compared.

In parallel at the first Hall sensor we measure the incoming pulses over 1 second (Timer2, Counter mode) and also the time (Timer4, Timer mode with 1µs time base) used for these pulses. 

Both - the speed and the acceleration - are transmitted via BLE every second.

After 30 seconds of measuring a pause of 40 seconds is given and the nRF is sent to energy saving mode; the acceleration sensor is configured to power save mode.

What we see is that after the around 20th work (30s) and pause (40s) cycle variations appear in the speed (of course in the work phase); also the number of reading-outs of the acceleration sensor decreases down to 0, but the cyclic transmission of the more and more instable speed and decreasing acceleration values is completely correct at the expected time points within the active 30 seconds...

The GPIOTE_IRQn is configured to APP_IRQ_PRIORITY_MID; the APP_TIMER_CONFIG_IRQ_PRIORITY was modified from 7 to 6 as described.

What can I check next? Or is there more to change in sdk_config.h?

Thanks in advance for your help.

Best regards,

matthK

Parents
  • Hi

    You mention that you are doing a lot in the interrupt. Have you done any measurements to verify how long time you are spending in the interrupt context? 

    For an application like this I think it would be interesting to use a scope or logic analyzer to look at the timing of the various activities going on in your system, which might give some indication what is happening when the system gets unreliable. 

    If you have any free pins on your board (or a LED pin or similar that can be temporarily re-purposed for something else) then you can toggle a pin when the interrupts start and stop in order to see how much time you spend on the interrupt processing, and where the interrupt happens in relation to other system events (such as the sensor interrupt and the hall sensor inputs). 

    It would also be interesting to see on the scope how the sensor readings are slowing down. 

    Best regards
    Torbjørn

  • Hello Torbjørn,

    thanks for your answer.

    I have connected a scope to a free pin and set it just at the beginning of the interrupt service function from the buffer full external interrupt from the accelerations sensor; just before leaving this ISF it is cleared. The ISF containing the reading out of the acceleration sensor buffer, calculation of the vector length and detection the biggest one takes 460 µs; this appears every 6.5 ms.

    I will connect a second pin which can be set and cleared at the start and end of the Hall sensor interrupt function for visualizing the system context.

    For the slowing down I have to think of a good trigger condition; this is perhaps not so easy...

    Best regards,

    matthK

  • Hi 

    That's a long interrupt handler for sure, looks like the GPIOTE interrupt lasts for more than 50ms? 

    From the graph it looks like the sensor buffer full interrupts are blocked by the GPIOTE interrupt, or are they not supposed to happen when you are in this state?

    If you are running such a long interrupt handler in the APP_IRQ_PRIORITY_MID priority it means you are blocking most other interrupts in the system, including callbacks from the SoftDevice informing you that something happened in the Bluetooth stack. It is quite likely that this will lead to issues with stability and performance. 

    You could try just to set the interrupt priority to LOWEST or LOW instead of MID and see if it works better. Ideally though I would recommend running such a long processing stage from main/thread context. 

    Best regards
    Torbjørn

  • Hi Torbjørn,

    I have made tests with LOW and LOWEST interrupt priority, but no change of behavior.

    But we found another point: When we ignore (see code below: added break statement) the PM_EVT_STORAGE_FULL event in the PeerManager  then it works perfect. 

    This is the code we have:

            case PM_EVT_STORAGE_FULL:
            {
               //  Run garbage collection on the flash.
            	break;
                err_code = fds_gc();
    
                uint32_t peerCount;
                peerCount = pm_peer_count();
                if (peerCount >= 8) // delete when 8 peers are connected
                {
                	pm_peer_id_t peer_id_to_delete;
    				err_code = pm_peer_ranks_get(NULL, NULL, &peer_id_to_delete, NULL);
    				if (err_code == NRF_SUCCESS)
    				{
    					NRF_LOG_INFO("Deleting lowest ranked peer (peer_id: %d)", peer_id_to_delete);
    					pm_peer_delete(peer_id_to_delete);
    				}
                }
    
                if (err_code == FDS_ERR_BUSY || err_code == FDS_ERR_NO_SPACE_IN_QUEUES)
                {
                    // Retry.
                }
                else
                {
                    APP_ERROR_CHECK(err_code);
                }
            } break;
    

    So we assume that the garbage collector fds_gc() cause the issue. 

    Should we keep it like that or replace by other code or execute at another time/place?

    Best regards,

    matthk

  • Hi 

    It makes sense that running this function in the middle of your sampling would cause issues. Deleting a page in flash could take as much as 80ms, and when running garbage collection multiple pages need to be erased at once. 

    It would make more sense to check the status of the fds library occasionally using the fds_stat(..) function, and running garbage collection at some opportune moment when there is not a lot of other activities going on (based on the number of dirty pages reported by fds_stat()). For instance it is quite common to schedule garbage collection after you get a disconnect on the Bluetooth connection, since at this point you know that there will be no Bluetooth activity. 

    It is a bit of a bad sign if garbage collection has to be run very often though. This implies that you have very frequent writes to flash, which will lead to significant wear over time. 

    To you know if the Bluetooth client enables or disables notifications often by writing to the CCCD?

    Best regards
    Torbjørn

  • Hello Torbjørn,

    thanks for this input; we will work on that in the next days and run a lot of tests - might be that we close this subject not before the end of this year...

    What I can say is that even when commenting out the garbage collector we had some variations in the speed and acceleration values; when reducing the IRQ priority for GPIOTE from APP_IRQ_PRIORITY_MID to APP_IRQ_PRIORITY_LOW the values became perfect. The GPIOTE-, SAADC- and APP_TIMER-IRQ priorities are configured to 6 in sdk_config.h

    Best regards,

    matthk

  • Hi

    It makes sense that lowering the priority of the GPIOTE interrupt would help. The 50ms interrupt in your diagram clearly appears to block the sensor interrupts, causing them to be delayed by more than 50ms. I am sure if you make another plot now it will look very different. 

    If you have any more questions later or have some test results to share just let me know. 

    Best regards
    Torbjørn

Reply
  • Hi

    It makes sense that lowering the priority of the GPIOTE interrupt would help. The 50ms interrupt in your diagram clearly appears to block the sensor interrupts, causing them to be delayed by more than 50ms. I am sure if you make another plot now it will look very different. 

    If you have any more questions later or have some test results to share just let me know. 

    Best regards
    Torbjørn

Children
No Data
Related