Is it possible to access a QSPI flash from nrf5340 as 'direct' memory access?

I have a custom board, with nrf5340 and a external QPSI flash (this is exactly like the nrf5340 DK).

I use the external flash for multiple partitions (mcuboot slots, FAT-FS file system, etc), and would like to use a partition as a "memory mapped file system" ie have blocks of data in the flash partition and allow code to access them as though it was data in the internal flash (ie just with a pointer in the 0x100000000 region). This lets me use this const data as though it was in my main application build but without having to use up space in the (limited) internal nrf5340 flash...

I can use the flash_open/read/write()  interface driver api to set up the partition data just fine.

However, when the code attempts to access the memory space directly (as a pointer access for example), it faults...

eg:

// the QPSI flash should be mapped at 0x10000000 and my partition starts at offset 0x780000 in my pm_static.yml

uint32_t* aptr = (uint32_t*)(0x10000000 + 0x780000); 

uint32_t val = *aptr;    // fault!

Your AI says there is no DTS or linker magic required to let the code read from this area, and I have checked that the flash initialisation priority (41 for the driver and 50 for flash) mean it should be initialised before my application starts running (priority 90).

The code also accesses the partition with flash_area_open() so that it can write to it via this api : could this cause an issue?

thanks!

Brian

Parents
  • not explicitly, but it has

    #define CONFIG_XIP 1
    in the autoconf.h si I guess yes!
    Also, I use this external flash for the nrf70 firmware which I think requires XIP to be enabled?
  • Does this mean that XIP must be enabled to allow this kind of read access? 

    I don't think it is, as if I enable it programmatically:

    nrf_qspi_nor_xip_enable(_ctx.flash_device, true);
    then all the QPSI flash accesses FAIL due to the ERRATA_159 workaround, which requires a specific HCLK (192MHz) and processor clock (64MHz) to access the QSPI...
    This implies that accessing the QSPI flash for XIP or data storage will not let me run at 128MHz???
    that shows such fast access?!
    What is the correct setup to get both performance and extended data storage from the QSPI?
  • BrianW said:
    Also, I use this external flash for the nrf70 firmware which I think requires XIP to be enabled?

    Been a while since I looked into that, but from my memory it depends on what mode you choose for the firmware patch. https://nrfconnectdocs.nordicsemi.com/ncs/latest/nrf/app_dev/device_guides/nrf70/fw_patches_ext_flash.html mentions two methods, one with xip and one with non-xip. I assume that you're using https://nrfconnectdocs.nordicsemi.com/ncs/latest/kconfig/index.html#SB_CONFIG_WIFI_PATCHES_EXT_FLASH_XIP since it enables it in autoconf.

    BrianW said:
    Does this mean that XIP must be enabled to allow this kind of read access? 

    That is at least my understanding of it. 

    The nRF5340 can read QSPI flash through a pointer at 0x10000000, but only while the QSPI peripheral is held in XIP mode. Errata 159 also requires the CPU clock (HCLK) to be 64 MHz for every QSPI transfer, and XIP fetches count. So with the workaround in place, you can't memory-map the flash and keep the CPU at 128 MHz at the same time. Without XIP, nrf_qspi_nor.c deactivates QSPI after every API call  qspi_release() calls nrfx_qspi_deactivate()). With PM it can also send the flash into deep power-down and uninit QSPI. A pointer read in that state gives a bus fault.

    I tried mirroring your setup with a nRF5340DK + 7002-EK shield, Wi-Fi fw patch in external flash and to read from the same range of addresses as you did. 

    Pointer reads into external flash work when all three of these are true:

    1. XIP is enabled with nrf_qspi_nor_xip_enable(dev, true) before the pointer read, and stays enabled while the pointer is used.
    2. The errata 43 workaround is turned back on in the application CMakeLists.txt.
    3. The CPU runs at 64 MHz while XIP is enabled.

    With these, pointer reads and flash_area_read() both returned correct data in every test, including after XIP was turned off again. For 128 MHz, also turn on the errata 159 workaround and never raise the clock while XIP is on.

    # Application CMakeLists.txt, after find_package(Zephyr) zephyr_compile_definitions(NRF53_ERRATA_43_ENABLE_WORKAROUND=1) zephyr_compile_definitions(NRF53_ERRATA_159_ENABLE_WORKAROUND=1) # only needed for 128 MHz

    nrf_qspi_nor_xip_enable(qspi_dev, true); uint32_t val = *(const volatile uint32_t *)(0x10000000 + 0x780000); nrf_qspi_nor_xip_enable(qspi_dev, false); /* when pointer access is done */


    Might be a bit hard to read the matrix below, but these are the configurations I tested on bench.

    TL;DR: XIP is required for pointer read and the order of enabling vs reading is important.

    I can share the heavily LLM assisted bench sample I used to validate if you wish to test it on your end as well. Happy to be challenged if you think I'm wrong on the XIP requirement + CLK frequency + order of enabling peripheral vs reading.

    Kind regards,
    Andreas

  • Thanks Andreas. 

    My main issue is then that I really don't want to have to run my CPU at 64MHz all the time just in case some of my code needs to access the external flash in XIP mode. Pretty much says using external flash for general data store is only for non-performance devices... 

    Using the flash_area_X() functions is ok for the CPU clock as these will do explicit qspi_acquire()/release() calls around their flash accesses, and these will change the hfclk192/hclk divisors on entry to work with the qspi bug 159, and can restore the hfclk back to the previous value on release() [this operation is my own patch, which was rejected by Nordic btw...]. Hence I can run at 128MHz and use flash_area without issues (just a little slower)

    But for XIP accesses there is no acquire/release handling, its just code accessing data pointers. So either the CPU clock needs to be specifically set to 64MHz all the time, or you can surround such code with calls to change the CPU clock... (effectively this is what calling xip_enable(drv, true) does for you). 

    I can hack this by identifing where in my code or libraries it uses the data I put in the flash, and wrapping that with xip_enable/disable... but its not great. For example, putting TLS certificates in external flash is nice as they are often 2-4kB these days, but when will the mbedtls library access them? who can tell?

    This also deals with the the other XIP issue : you don't have a way to power off/on the external flash when you're not accessing it... 

    TLDR; XIP to external flash is only useful if you don't care about performance and battery life, unless you can exlicitly determine when your XIP accesses are happening.

    For background, the reason I wanted to push data out to external flash is that when trying to run the wifi in STA mode. The zephyr wifi stack and particularly wpa-suppliant/mbedtls use so much flash space its a big struggle to find space for the application....

    Now if your LLM can generate a project where the whole of wpa_suppliant/ related mbedtls dependancies can be placed in a XIP segment and enabled when run at AP authentication time and then disabled once its done, that would be very useful! Thats around 250kB of internal flash saving right there I reckon....

Reply
  • Thanks Andreas. 

    My main issue is then that I really don't want to have to run my CPU at 64MHz all the time just in case some of my code needs to access the external flash in XIP mode. Pretty much says using external flash for general data store is only for non-performance devices... 

    Using the flash_area_X() functions is ok for the CPU clock as these will do explicit qspi_acquire()/release() calls around their flash accesses, and these will change the hfclk192/hclk divisors on entry to work with the qspi bug 159, and can restore the hfclk back to the previous value on release() [this operation is my own patch, which was rejected by Nordic btw...]. Hence I can run at 128MHz and use flash_area without issues (just a little slower)

    But for XIP accesses there is no acquire/release handling, its just code accessing data pointers. So either the CPU clock needs to be specifically set to 64MHz all the time, or you can surround such code with calls to change the CPU clock... (effectively this is what calling xip_enable(drv, true) does for you). 

    I can hack this by identifing where in my code or libraries it uses the data I put in the flash, and wrapping that with xip_enable/disable... but its not great. For example, putting TLS certificates in external flash is nice as they are often 2-4kB these days, but when will the mbedtls library access them? who can tell?

    This also deals with the the other XIP issue : you don't have a way to power off/on the external flash when you're not accessing it... 

    TLDR; XIP to external flash is only useful if you don't care about performance and battery life, unless you can exlicitly determine when your XIP accesses are happening.

    For background, the reason I wanted to push data out to external flash is that when trying to run the wifi in STA mode. The zephyr wifi stack and particularly wpa-suppliant/mbedtls use so much flash space its a big struggle to find space for the application....

    Now if your LLM can generate a project where the whole of wpa_suppliant/ related mbedtls dependancies can be placed in a XIP segment and enabled when run at AP authentication time and then disabled once its done, that would be very useful! Thats around 250kB of internal flash saving right there I reckon....

Children
  • BrianW said:
    My main issue is then that I really don't want to have to run my CPU at 64MHz all the time just in case some of my code needs to access the external flash in XIP mode. Pretty much says using external flash for general data store is only for non-performance devices... 

    Very understandable. I don't disagree with the sentiment. If I remember correct from ~3-4 years ago, there was a general recommendation (I can't remember if it was "official" statement or just an engineering recommendation) w.r.t the XIP QSPI feature for the nRF5340, but to summarize: The recommendation back then was that if you're going to execute code from external flash with XIP over QSPI it was recommended to

    1. Not have any timing sensitive code/interrupt based code executed from external flash due to execution speed + time to init / deinit the qspi
    2. Not have the feature on a battery powered device with a small battery due to increased current consumption. The feature is best suited for MAINS powered devices due to cost with both having the external flash powered and to init/de-init the QSPI

    I assume that you're already very committed to this project + BOM and socs, but just in case this is still in early evaluation from your end. Since it seems to me that memory might be a limitation. Is it an option to swap the host MCU to a nRF54LM20 to avoid having to use the XIP over QSPI, or has that ship sailed for this product/does that still not fit within your flash/RRAM requirements?

    BrianW said:
    Now if your LLM can generate a project where the whole of wpa_suppliant/ related mbedtls dependancies can be placed in a XIP segment and enabled when run at AP authentication time and then disabled once its done, that would be very useful! Thats around 250kB of internal flash saving right there I reckon....

    There should be built in support to have the WPA supplicant + related mbetls in a XIP segment, but since supplicant keeps running after it connects we can't power of the XIP partition without any risk. There's a rekey interval that we will need ensure that we have the external flash powered on, as I'm sure you're aware of.

    The WPA supplicant handles group and pairwise rekey (EAPOL frames), deauth and disconnect events, roaming and background scans, SA query, and wifi status API calls. Its event loop thread is also blocked inside hostap code. Any of these with XIP off gives a bus fault. 

    But that said, it shouldn't be impossible, it's just a bit risky. I'll do an over the weekend test with a gated implementation (i.e turning on the XIP when we need to use the supplicant) to see if it works in my setup. The rekey interval is relatively high on the access point I'm using to test, so I won't get any significant results in by today.

    Before those results are in, here's the results of relocating the hostap in xip on my setup. The Wifi FW patch is in the external flash using SB_CONFIG_WIFI_PATCHES_EXT_FLASH_STORE, and not SB_CONFIG_WIFI_PATCHES_EXT_FLASH_XIP (check which you have in the generated config).

    Pass Build Internal flash Boots OK Ping replies Connect time Idle current
    0
    Control
    609 308 B
    5/5
    50/50
    5.21–5.26 s
    1.392 mA
    1
    hostap code in XIP
    486 800 B
    5/5
    50/50
    5.19–5.24 s
    4.058 mA
    2
    XIP on, nothing relocated
    609 416 B
    5/5
    50/50
    5.18–5.31 s
    4.063 mA

    Current is as you see relatively high, and it might be higher as I don't measure the external flash usage here, only VDD_nrf through a PPK2 and as of now the external flash is on VDD_PER.

    Kind regards,
    Andreas

  • Not have the feature on a battery powered device with a small battery due to increased current consumption.

    Ah... thats exactly my use case

    Since it seems to me that memory might be a limitation. Is it an option to swap the host MCU to a nRF54LM20 to avoid having to use the XIP over QSPI, or has that ship sailed for this product/does that still not fit within your flash/RRAM requirements?

    we are already in production and deployment. Currently I am trying to squeeze more functionality in based on customer projects - and trying to get the wifi enterprise stuff in there is a problem....

    A bit of speed hit is managable, but a battery hit will be a problem - I'm already struggling with my sleep current on the board (up around 5mA when it should be <1mA). That said, I'm pretty happy with the current used in operation with the wifi connected running MQTT.... 

    The WPA supplicant handles group and pairwise rekey (EAPOL frames), deauth and disconnect events, roaming and background scans, SA query, and wifi status API calls. Its event loop thread is also blocked inside hostap code. Any of these with XIP off gives a bus fault. 

    ah, true. for roaming, I understood it doesn't currently support any fast roaming (802.11r), so I've been handling this from the application by checking wifi_status and doing a reconnect as fast as I can... Since I manage this from the app, I can explicitly turn on/off XIP...

    And managing the XIP on/off from the application is working for my 'memory map data files' operations (for fonts, I have't tried it for https TLS certificates yet....)

    Before those results are in, here's the results of relocating the hostap in xip on my setup.

    Interesting, yes. I use SB_CONFIG_WIFI_PATCHES_EXT_FLASH_STORE (hence my XIP wasn't on by default...). 

    I guess the current being nearly 4x higher with XIP on is just beacause it leaves the HCLK192 at DIV1 and not DIV4? Or are you running PM and getting the enter_dpd/exit_dpd benefit in non-XIP mode?

    I'm not running PM - maybe this is causing some of my high 'sleep' current...

    thanks for testing all this!

Related