Repository navigation
[FEATURE]: Provide an obvious way to access underlying data arrays within a Figure object (rather than the base64-encoded versions) #5124
Description
Activity
This may not solve your issue, however I had a similar issue in a client application, rendering the json object using the plotly.js library.
It turned out I was using an outdated plotly javascript library, which in versions prior to v2.28 did not know about bdata (base64 encoded). Once I upgraded to the latest version of plotly.js it worked fine.
@marthacryan suggests adding a base64: bool argument to plotly.io.to_json (similar to pretty: bool) that would control whether data is base64'd or not.
For some context on implementation: the BaseFigure class does the conversion to base64 in the to_dict function here:
plotly.py/plotly/basedatatypes.py
Line 3344 in 1ec864b
| convert_to_base64(res) |
to_jsoncallsvalidate_coerce_fig_to_dictherevalidate_coerce_fig_to_dictcallsBaseFigure.to_dicthere.
So, we'd have to add a flag to plotly.io.to_json and validate_coerce_fig_to_dict and BaseFigure.to_dict.
I updated this issue title to cover a broader feature request: given a Figure object, plotly.py should provide a simple way to access the underlying data values without base64 encoding.
Context: Plotly.py consists of two layers:
- The outer Python layer, which contains the Python API (
plotly.express,plotly.graph_objects, etc.) - The underlying JavaScript layer (Plotly.js), which runs in the browser (or in Kaleido) and creates the actual chart visuals
JSON is the communication between these two layers. To create a chart, Plotly.py generates a JSON-formatted string containing the full chart spec (data, layout, colors, font options, etc.), and sends the JSON string to the browser, where Plotly.js interprets it and renders the chart.
In versions of Plotly.py prior to 6.0.0, all data arrays (such as the "x" and "y" arrays for a scatter plot) were represented in JSON as simple lists of values:
"y": [1,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57,58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91,92,93,94,95,96,97,98,99,100]Plotly.py 6.0.0 introduced base64 encoding of typed arrays (implemented in #4470 and plotly/plotly.js#5230).
When a Plotly.py figure contains an explictly-typed data array (such as a numpy array or Pandas series), the array is represented as a dictionary with a "bdata" value containing the array data encoded as a base64 string, and a "dtype" value specifying the data type:
"y": {
"dtype": "i1",
"bdata": "AQIDBAUGBwgJCgsMDQ4PEBESExQVFhcYGRobHB0eHyAhIiMkJSYnKCkqKywtLi8wMTIzNDU2Nzg5Ojs8PT4/QEFCQ0RFRkdISUpLTE1OT1BRUlNUVVZXWFlaW1xdXl9gYWJjZA=="
}Plotly.js decodes the 'bdata' string and recovers the original data.
This change has a few advantages:
- Faster JSON serialization (3-10x faster from a quick benchmark I did just now)
- Somewhat shorter JSON string (40-50% shorter for floats, although less difference after compression)
- Ensures Plotly.js receives numbers with the exact same precision as the ones specified in Plotly.py
However, there's also one big disadvantage: encoded values are not human-readable (or computer-readable) without decoding. So the changes in #4470 broke all workflows that involve inspecting the figure JSON to view or extract data values.
Specifically, the changes in #4470 mean that some of the main methods for introspecting inside a figure object now contain b64-encoded values, including fig.to_dict() and fig.to_json().
We should expose at least one method which lets you retrieve the whole figure object without base64-encoding the arrays. I'm reviewing the existing methods right now and the situation is a bit messy, so there might be some additional cleanup that can be done as part of the same effort.
Hi, I have some some code here, as below (edited to remove extra logs). I'm converting a plotly figure into json, and then loading it with json loads. You would expect that to preserve the original data.
However, as you can see, the "y" data (which had previously been an array of large numbers) are wrongly converted into this "bdata" thing instead.
I was able to solve this and get the data in the format I wanted by going from 6.0.1 -> 5.24.1, which was the last stable build I was using. I think this is a pretty fundamental issue, so let me know if there's anything else I can share.